Skip to content

Project Management Mastery / Chapter 15

Estimate Honestly Under Uncertainty

An estimate is a decision instrument, not a number factory. This chapter separates estimate, target, commitment, and forecast, follows the cone of uncertainty, names the pressures that bend numbers low, and builds an honest range with its arithmetic, its pockets of contingency and management reserve, and its basis of estimate — so the project hands the sponsor a range instead of a precise fiction.

Chapter 15: Estimate Honestly Under Uncertainty

The number the city asked for

It is month seventeen at BlueLine, and the city wants one number. Marta Reyes, the transport authority’s finance officer, has written to the consortium for the second-half authorization: what will the rest of the corridor cost, and will it open when the mayor promised? “A figure,” she writes, “not a range. The council needs a figure to appropriate against.” The corridor is seven months from its target opening, and about 1,150 million units of the 2,400 million-unit capital envelope are already committed. The remaining ask is the 1,250 million units to finish the corridor as engineered, plus the 140 million-unit mitigation line the business case carries: a forward base estimate of about 1,390 million units. Behind that number sits a mix of engineered civils, unproven systems integration, and an access standard whose final scope has not been signed.

Daniel Osei, the consortium’s finance director, has the authorization pack on his desk. The first page contains the answer the city is asking for: “remaining forward cost, 1,486.7 million units.” A junior analyst produced it that morning by adding the latest forecasts from the four delivery partners. The number carries three decimal places of authority no one in the building believes. Daniel’s own team ran a different exercise: a range, built from the same inputs but with the uncertainty left in. The forward base estimate is about 1,390 million units, and the 80-percent range runs from roughly 1,280 to 1,500. Below 1,280, the eastern flood-plain segment has beaten its drainage approval, the ticketing integration has gone cleanly, and the access standard has stayed inside its draft scope. Above 1,500, the flood plain, the integration, and the access standard have all gone badly, which the risk analysis considers entirely plausible.

The meeting that follows is short and uncomfortable. The contractor’s commercial director says the city needs a number, and the precise number is the number. Miguel Alvarez says the transport model cares about two endpoints, the cost and the date, and it can absorb a range better than it can absorb a surprise. Grace Njoroge, whose federation co-authored the access standard, says the range is the only honest answer, because the access standard’s scope is a decision, not a discovery. Marta’s deputy, on the call from the mayor’s office, says what everyone knew he would say: “If we go to council with ‘somewhere between 1.28 and 1.50 billion,’ the opposition will eat us alive.”

Lena Voss, the project leader, has heard enough. “Then we go to council with the truth and we teach them how to read it,” she says. “The moment we hand them one number, we are betting the whole authorization on a fiction. The range is not a failure of estimating. It is the estimate.”

This chapter is about that sentence. Estimation is the Delivery lens at its most exposed: the number is the point where the project’s uncertainty becomes public, and the instinct of every room is to hide the uncertainty behind arithmetic. The mastery the chapter teaches is the opposite. An honest estimate is a decision instrument: it tells the sponsor what the work will cost and when it will finish, how sure anyone is, what would have to be true for the number to move, and where the money to absorb the surprise is held. A dishonest estimate, precise, brave, and wrong, is the most expensive document a project produces, because every later decision is built on it. Chapter 14 built the structure the estimate lands on, the work packages, the dictionary, the seams. This chapter gives the structure numbers that tell the truth about themselves.

Four words that look like one number

The first discipline is vocabulary, because the room at BlueLine was arguing about one number when it was really using four words as if they were one.

An estimate is an analytical projection: the best current description of what a quantity will be, based on the available evidence, with its uncertainty stated. It is produced by the people who can examine the evidence: the engineers, the analysts, the people who will do the work. It is revised when the evidence changes.

A target is a desired value: what the sponsor wants to achieve, the cost ceiling or the date the organization has chosen to aim at. A target is an intention, not an analysis. It belongs to the sponsor, it may be ambitious, and it is exactly as honest as its owner is willing to say that it is an ambition.

A commitment is an accepted obligation: a promise, made by a person with the authority to bind the project, to deliver within stated limits. A commitment is made after the estimate, the risk, and the capacity have been weighed. It is the one number that has consequences, and the one number that changes only through the change control that chapter 39 will build.

A forecast is the current evidence-based expectation: the number the project would defend today, given everything it now knows. The forecast updates as evidence arrives, by design, and its history is a learning record.

The failure is not that the four words exist. The failure is that they are used interchangeably, because each one has a different owner, a different update rule, and a different consequence. Consider the sentence every project says at least once: “the corridor opens in month twenty-four.” As a target, it is the mayor’s promise to the council, ambitious and public. As an estimate, it is the analysts’ best current projection of when the corridor can actually open, probably different from the target. As a commitment, it is the obligation Lena accepted on behalf of the consortium, binding under the governance of chapter 8. As a forecast, it is the number the project would defend at this month’s steering committee, and it has probably moved since the estimate was written. One sentence, four meanings, four owners, and the room cannot argue honestly about any of them until it says which one it is using.

The current Scrum Guide (November 2020, by Ken Schwaber and Jeff Sutherland) encodes the same discipline in software terms. The work selected into a Sprint is discussed as a forecast, and the Guide grounds the developers’ confidence in that forecast in their past performance, their upcoming capacity, and their definition of done. The Product Goal, the Sprint Goal, and the Definition of Done are the commitments each artifact carries. The vocabulary exists because the confusion is expensive: a Sprint forecast is evidence-based and updates every Sprint, while a Sprint Goal holds for the Sprint and changes only when it becomes obsolete. The discipline transfers to any project, predictive or adaptive: label the number before debating it.

The practical move is a small one, and it belongs in every planning meeting: when someone proposes a number, ask which kind it is. “Is that the estimate, the target, the commitment, or the forecast?” The question forces the room to name the obligation. A target that is being read as an estimate is the number that must be hit wearing an analyst’s coat; an estimate being read as a commitment is a promise nobody authorized; a forecast being read as a target is the project being blamed for the sponsor’s ambition. None of those confusions is resolvable until the label is spoken.

The shape of honest uncertainty

The second discipline is knowing how wide honest numbers are. In 1981, Barry Boehm’s Software Engineering Economics charted the range of effort estimates across the life of a software project and found that early estimates could be off by a factor of about four in either direction: the actual effort might be a quarter of the first estimate, or four times it. Boehm called the pattern the funnel curve; Steve McConnell gave it the name it is known by, the cone of uncertainty, in Software Project Survival Guide (Microsoft Press, 1997), and built the practical treatment in Software Estimation (Microsoft Press, 2006). The pattern predates software: the founders of the American Association of Cost Engineers, the body now known as AACE International, published cost-estimate classification work in the chemical industry in 1958. The earliest base estimates of a plant’s cost were often exceeded by as much as 100 percent, or underrun by as much as 50 percent.

Chapter 13 used the cone to choose the life cycle: put the expensive discovery early, where the cone is wide, and reserve commitments for where it is narrow. This chapter uses the same cone for a different job: it is the estimate’s own honesty instrument, the shape that says how much any number at this moment can be trusted. The cone’s width is information. An estimate without a range is an estimate with its uncertainty deleted, and the reader of that estimate, the sponsor, the council, the public, will supply a range of their own, usually narrower and always wrong.

Two properties of the cone matter here, and both are routinely misunderstood.

The first is that the cone narrows only when uncertainty is retired by decisions and evidence, never by the passage of time. A scope decision narrows it: the access standard’s final scope, the ticketing release boundary, the flood-plain drainage approval. A prototype narrows it: the integration spike that proves the vendor’s system can talk to the legacy scheduling platform. An accepted test narrows it: the clinic wave that proves the workflow under real workload. But a project that simply waits, that holds its meetings and updates its dashboard while making no decisions and testing nothing, keeps its month-one cone at month ten. The calendar does not retire uncertainty. The cone is a map of knowledge, and knowledge grows from decisions, not from the date.

The second property is the one the city’s question exposed: the cone does not forgive precision. A number with three decimal places carries no more knowledge than a number with none, and it can carry less, because the decimals suggest a measurement where there was only arithmetic. The precise figure the junior analyst produced, 1,486.7 million units, was a fiction not because the arithmetic was wrong but because the inputs were ranges and the output was a point. The arithmetic can be perfect and the estimate still false, when the uncertainty was deleted somewhere between the inputs and the sum.

The estimate maturity idea follows from the cone: an early estimate is an order-of-magnitude number, useful for deciding whether to investigate; a mid-project estimate is a budget number, useful once the requirements are understood; a late estimate is a definitive number, useful when the design is complete and the remaining uncertainty is real but small. Cost engineering practice formalized this intuition: AACE International’s recommended practices classify estimates from class five, for concept work with the widest accuracy range, to class one, for near-complete engineering with the narrowest, with each class tied to the maturity of the project definition. The naming does not matter; the honesty does. A project that presents a class-five estimate with class-one precision has lied without writing a false number.

The pressures that bend the number

The third discipline is knowing why estimates drift, because the drift is systematic and it is not primarily about arithmetic. Three forces bend project numbers low, and the honest estimator must name all three before producing a range.

The first force is the planning fallacy, the pattern Kahneman and Tversky first described in their 1979 work on intuitive prediction: people make plans by imagining the future of the specific task, simulating the steps as they hope they will go, and they forget the distribution of how such tasks have actually gone. The classic demonstration is Roger Buehler, Dale Griffin, and Michael Ross’s 1994 study of students predicting their senior theses. The students predicted an average of 33.9 days and completed in an average of 55.5; their own worst-case predictions averaged only 48.6 days. The estimates were not dishonest; they were generated inside a simulation that left out the evidence. Dan Lovallo and Daniel Kahneman, writing in Harvard Business Review in 2003, expanded the definition: the planning fallacy is the tendency to underestimate the time, cost, and risk of future actions while overestimating their benefits. The inside view, the planner staring at the task and simulating the plan, is the engine of the bias.

The second force is anchoring, the phenomenon Tversky and Kahneman demonstrated in their 1974 heuristics and biases research: judgment is pulled toward whatever number was mentioned first, even when the anchor is transparently uninformative. In the original demonstrations, participants who saw a rigged roulette wheel stop at 10 guessed that about 25 percent of African nations were in the United Nations, while those who saw it stop at 65 guessed about 45 percent. Later work reproduced the pull with absurd anchors: students who were asked whether Gandhi died before age 9 or after age 140, then asked for their best guess, gave average answers of about 50 and 67 respectively. In an estimating room, the anchor is whatever number is spoken first, and it is usually spoken by the most senior person in the room. The budget number from last year, the sponsor’s target, the first engineer’s offhand guess, all anchor the conversation, and every subsequent estimate is an adjustment from the anchor rather than an independent judgment. The adjustment is never enough. The room that estimates the ticketing integration will end up near the first number mentioned, not near the evidence.

The third force is strategic misrepresentation, and it is the one that cannot be corrected by better psychology. Bent Flyvbjerg, Mette Holm, and Søren Buhl studied 258 transportation infrastructure projects worth about 90 billion dollars and found that costs were underestimated in almost 9 out of 10 projects, with actual costs on average 28 percent higher than estimated, and the underestimation varied by project type: 44.7 percent on average for rail, 33.8 percent for bridges and tunnels, and 20.4 percent for roads. Their conclusion, published in the Journal of the American Planning Association in 2002, is blunt: the underestimation cannot be explained by error, and is best explained by strategic misrepresentation, that is, lying. The mechanism is familiar to anyone who has sat in an authorization meeting: the project that asks for less than it needs gets approved, the overrun is discovered later, and by then the money is committed. The sunk cost argument, chapter 7’s hardest lesson, carries the project forward. Optimism bias and strategic misrepresentation are different diseases with similar symptoms. The sincere optimizer under-simulates the future; the strategic misrepresenter under-prices the present, because the project’s survival depends on the number. The corrections differ accordingly: the optimizer needs the outside view, the subject of a later section; the misrepresenter needs consequences, independent review, and a governance that does not reward the low number. The title of the 2002 paper, Error or Lie?, is the diagnostic question every estimate must answer for itself.

There is a fourth pressure, and it is the one the estimator creates alone, in silence: hidden contingency, the padding added to an estimate because the honest range was not presentable. The analyst who knows the 80-percent range but is told to produce a point adds 15 percent quietly, so that the point lands somewhere defensible. The padding is not declared, not budgeted, not managed, and it is consumed silently as the project proceeds, so the project looks healthy while the padding is spent. The overrun arrives with the padding gone and nobody able to say when it was used. The sandbag is the failure pattern of section after section, and it deserves its own name because it is produced by the same instinct that produces honesty: the estimator wanted to be truthful, found the truth unpresentable, and smuggled it in as a number. The repair is not to forbid the padding; it is to make the range presentable, which is the work of the rest of this chapter.

Five ways to make a number

The methods are where the craft lives, and each method exists to improve a particular decision. The discipline is to choose the method for the question, to record which method was used, and to know what each method is bad at.

Analogous estimation starts from a completed project that resembles this one, and scales its outcome. The corridor’s early numbers came from the city’s portfolio: per lane-kilometre costs from comparable bus corridors, adjusted for the flood plain and the access standard. The decision an analogous estimate improves is “is there a genuinely comparable outcome?” Its minimum viable form is one comparable project, named, with the differences stated and priced. Its failure is the comparable that is not comparable: the corridor in the other city with different geology, a different procurement regime, a different regulatory environment, quoted as if it were the same work. The analogous estimate is the fastest method and the easiest to fake, because the comparison is always partial, and the honest version says which parts of the comparison are solid and which are hopes.

Parametric estimation builds a model: cost as a function of a few drivers, the cost per lane-kilometre, the cost per function point, the cost per bed, the cost per release. In software, Barry Boehm’s COCOMO (Constructive Cost Model) line of models, first published in 1981, estimates effort from size and a set of cost drivers, and parametric estimation is the standard tool of cost engineering where historical data is rich. The decision a parametric estimate improves is “what actually drives the cost, and do we have the data to price the driver?” Its failure is invisible: the model’s coefficients come from data, the data comes from somewhere, and the fit rarely survives contact with a novel project. The parametric estimate inherits the quality of its data, and its uncertainty includes the model’s uncertainty, which the model itself cannot report.

Bottom-up estimation builds the number from the work packages chapter 14 taught: estimate each package, sum the packages, and let the dictionary keep the views honest. It is the most defensible method when the scope is well defined, because the estimate is then an audit trail: every unit is attached to an owned package with an acceptance. It is the most expensive method when the scope is not defined, because the team pays the full decomposition cost for fiction. It has one structural trap, and it is the chapter 14 trap wearing a financial costume. The sum is only as good as the packages that were included; the disappearing 100 percent reappears as the missing integration node, the missing rehearsal, the missing handover evidence — the work that lives in no contract and no estimate. The bottom-up estimate must be reconciled against the same completeness review as the decomposition, or the precision of the sum will certify the absence of the work it forgot.

Three-point estimation is the method that forces a range. It came from the Program Evaluation and Review Technique, PERT, developed for the United States Navy’s Polaris missile program by the Navy’s Special Projects Office with Lockheed and Booz Allen Hamilton and made public in 1958. PERT asks for three numbers per activity: the optimistic duration, the most likely, and the pessimistic, and computes an expected value as a weighted average, (optimistic + 4 times most likely + pessimistic) divided by 6, with a rough standard deviation of (pessimistic minus optimistic) divided by 6. The decision a three-point estimate improves is “how wide is the range, honestly?” Its minimum viable form is three numbers and the arithmetic, with the pessimistic as a genuine tail. Its failure is theater: the three numbers are produced to justify a desired answer, the most likely tuned to the target, the pessimistic trimmed to stay presentable, and the resulting range is as fictional as a single point. The method is honest only when the three numbers come from evidence and the pessimistic is allowed to hurt.

Structured expert judgment is the family of methods that pools independent estimates on purpose. The Delphi method, developed at the RAND Corporation by Olaf Helmer and Norman Dalkey in the 1950s and documented in a 1963 Management Science paper, collects anonymous estimates from a panel, summarizes them, and iterates toward convergence without letting the panel anchor on each other’s numbers. Wideband Delphi, originated by Barry Boehm and John Farquhar and popularized in Boehm’s 1981 book, adds group discussion between rounds. Planning poker, first defined and named by James Grenning in 2002 and popularized by Mike Cohn’s Agile Estimating and Planning (2005), is the same mechanism made fast: estimators reveal their numbers simultaneously, the outliers explain their reasoning, and the group converges or deliberately disagrees. The decision structured judgment improves is “how do we pool independent judgment without the first number winning?” The mechanism against anchoring is the simultaneous reveal, and the mechanism fails when the reveal is rushed, the explanation is skipped, or the loudest voice re-anchors the room anyway. The honest version of expert judgment produces a range with names attached: here is why the civil engineer sees 9 weeks, and here is why the systems lead sees 17.

The adaptive family adds a sixth way of working, and it is the one that changes the question. Relative estimation sizes work in units that compare items to each other, story points, rather than in time. The forecast is then produced from observed throughput: how many points the team has actually completed per week, and how long completed work has spent in the system. The mathematics underneath is Little’s Law, the queuing result John Little proved in 1961, which ties the work in progress in a system to the arrival rate and the time items spend inside it. Chapter 17 states the law with its assumptions and works its arithmetic; Daniel Vacanti’s Actionable Agile Metrics for Predictability (2015) is the practitioner treatment that built forecasting discipline on the foundation. The decision this way of working improves is the one KijaniPay kept asking: “when will it actually be done, from evidence?” The answer comes from history, not from hope. The failure is velocity theater: the team that re-sizes the points to fit the date, changes its definition of done and pretends the scale held, or games its work-in-progress limits to manufacture a number, produces a forecast as fictional as any precise liar’s. The throughput forecast is only as honest as the data and the discipline behind it; the principle here is that the adaptive family replaces the estimate of the future with the measurement of the past.

Every method produces the same artifact, and the artifact is the chapter’s quiet workhorse: the number plus its basis, the record of method, data, assumptions, and width. Without the basis, the number is an orphan; with it, the number can be examined, challenged, and improved, which is exactly what an estimate is for.

The arithmetic of a range

The numbers matter here, and they are simple enough to check by hand.

Take the ticketing integration at BlueLine, the work package whose seam chapter 14 found. The team produces three numbers: optimistic, 5 weeks; most likely, 9; pessimistic, 17. The PERT expected value is (5 + 4 times 9 + 17) divided by 6, which is 58 divided by 6, about 9.7 weeks. The rough standard deviation is (17 minus 5) divided by 6, 2 weeks, which is PERT’s convention: the optimistic and pessimistic values are treated as sitting about three standard deviations from the mean, which is why the spread is divided by six. Two numbers the room might have quoted instead are both wrong in instructive ways. The most likely value, 9 weeks, is what the team actually expects on a normal week, and it is below the expected value because the distribution is asymmetric: pessimistic outcomes are further from normal than optimistic ones. The naive midpoint, 11 weeks, the average of 5 and 17, treats the two tails as equally likely, which they are not. The expected value, 9.7 weeks, is the honest center of the distribution. If we approximate the shape as normal, the convention says about two-thirds of outcomes fall within one standard deviation, roughly 7.7 to 11.7 weeks. The chance of exceeding 11 weeks is only about a quarter, since 11 sits about two-thirds of a standard deviation above the mean.

Then the sensitivity, because the estimate is a bundle of assumptions and each assumption has a price. Suppose the pessimistic case is really 22 weeks, because the team has not scoped the legacy scheduling interface and suspects it is larger than the 17-week figure assumed. The expected value moves to (5 + 36 + 22) divided by 6, about 10.5 weeks, and the spread widens from 12 weeks to 17. One assumption, five weeks of movement in the pessimistic case, almost a week of movement in the expected value. That is why the assumptions log exists, and why the confidence statement must name the swing assumptions: the number is only as stable as its assumptions, and the honest estimate says which assumptions, if wrong, would move it.

The expected value also deserves a direct treatment, because it is a decision aid, not a prediction. The consortium has two commercial options for the integration. Option A is the proven vendor at a fixed cost of 6.0 million units with a narrow range. Option B is the newer vendor with incentive pricing: an 80 percent chance the integration costs 4.0 million units, and a 20 percent chance it costs 14.0, because the newer vendor has failed on a similar interface before. The expected value of B is 0.8 times 4.0 plus 0.2 times 14.0, which is 3.2 plus 2.8, exactly 6.0 million units, the same as A. The two options have the same expected value and completely different shapes: A is narrow and certain, B is bimodal, cheap most of the time and expensive one time in five. The decision between them is not decided by the arithmetic; it is decided by the corridor’s capacity to absorb the 14.0 tail against the political cost of a failed integration, by the risk appetite the governance of chapter 8 has stated, and by the value of the downside option. The expected value is the honest comparison surface, and the shape is the decision. Anyone who reports only “6.0 million units either way” has reported the arithmetic and deleted the decision.

The third arithmetic is the confidence statement, and it is the sentence that makes an estimate checkable. At BlueLine it reads: “The forward base estimate is about 1,390 million units. The risk analysis puts the 80-percent range at roughly 1,280 to 1,500 million units. The three largest assumptions are the eastern flood-plain drainage approval, the ticketing integration outcome, and the final scope of the access standard; each is owned, and each has a trigger date.” The 80-percent range is a property of the model and its assumptions, not a promise about the world: if the model is right, four times out of five the final number falls inside the range. The range can still be wrong. What the statement does is make the estimate defeatable in advance: the reader can check the assumptions, watch the triggers, and see the range narrow as the uncertainty retires. That is chapter 7’s sensitivity discipline, once the project is inside the work.

Two pockets, two owners

The range converts into money through two pockets, and the discipline is keeping the pockets distinct, because their confusion is a common corruption.

Contingency is the pocket for known unknowns: the risks that can be named, the flood-plain approval, the integration outcome, the access-standard scope, the things the risk register of chapter 22 will carry. Contingency is sized from the range, it is held by the project, it is part of the approved budget, and it is released against named risks by decision, with a trigger and evidence. Its purpose is not to avoid being spent; its purpose is to be spent when the named risk materializes. A contingency that is never spent is a sign either of luck or of an estimate that was wider than it needed to be, and both are findings, not complaints.

Management reserve is the pocket for unknown unknowns: the surprises that cannot be named, the things nobody thought to list, the flood that is not the flood-plain approval. Management reserve sits above the project, at the portfolio or the sponsor’s level, it is not part of the project’s baseline, and it is released by governance, not by the project team. The distinction matters because the incentives differ: the project team is accountable for contingency, so it is spent against evidence; no one is accountable for reserve until the governance releases it, so it stays until the surprise is real.

The BlueLine numbers make the pockets visible. The forward base estimate is 1,390 million units. The risk analysis puts the 80th percentile at about 1,500, so the contingency is about 110 million units, roughly 8 percent of the forward base, held by the consortium and managed against the named risks. Beyond 1,500, the modeling runs to a 1,570 ceiling, and the corridor’s steering committee has decided that the band from 1,500 to 1,570 is the city’s reserve: the city holds it, the project does not, and it is released only through re-authorization, with the council’s visibility. The structure is deliberate: the project manages the risk it can name, the city prices the risk it cannot, and neither pocket is invisible.

The failure pattern is the merged pocket. The project that holds its contingency as a secret margin, spent quietly on ordinary overruns, has turned contingency into the sandbag: the named risks arrive with the pocket empty and no one able to say where the money went. The project that treats its management reserve as its own, drawing it down without governance, has moved the unknown-unknowns onto its own books and removed the one check that makes reserve honest. And the project that reports “we are under budget” while its contingency drains is reporting the empty pocket as health, which is the precise liar in a new costume. Chapter 18 will build the cost control machinery that tracks these pockets monthly; the principle is established here: the pockets exist, they are labeled, and their spending is a decision with evidence.

The outside view

There is a correction to the planning fallacy that deserves its own section, because it is the single most reliable debias an estimator has: the outside view.

Lovallo and Kahneman, in the 2003 Harvard Business Review article, named it reference class forecasting: instead of estimating the specific project from the inside, estimate it from the distribution of comparable completed projects. Ask not “how will this corridor go?” but “how have corridors like this actually gone?” The reference class does the work the inside view cannot: it contains the failures the current team has never experienced, the flood-plain approvals that took two years, the integrations that doubled, the access standards that grew. The inside view simulates a future; the outside view remembers a past, and the past is the better predictor.

The method became official in UK transport planning: in June 2004, the Department for Transport published Procedures for Dealing with Optimism Bias in Transport Planning, guidance developed with Bent Flyvbjerg and the consultancy COWI, which required reference-class uplifts to be applied to transport project estimates before authorization. The guidance’s first practical test is a documented lesson, and it is worth telling straight.

In October 2004, Ove Arup and Partners Scotland reviewed the Edinburgh Tram Line 2 business case. The project was then forecast to cost a total of 320 million pounds, including 64 million pounds of contingency, about 25 percent on the base estimate. Using the new reference-class guidance, the reviewers calculated that the 50th percentile of total capital cost, the point with a 50 percent chance of staying within budget, was 357 million pounds, requiring 40 percent contingency, and the 80th percentile was 400 million pounds, requiring 57 percent contingency. The reviewers explicitly warned that even their reference-class forecasts were likely too low, because the uplifts should have been applied at the decision-to-build point and the project had not reached it. The tram line opened in May 2014, three years late, with a final outturn cost of 776 million pounds, about 628 million in 2004 prices, nearly double the original forecast. The reference-class numbers were wrong, and they were closer than the business case: 357 million was nearer the truth than 320, and the reviewers’ warning that even their number was too low was the most accurate sentence in the file.

The lesson is not that reference classes are magic; it is that the outside view is a check, not a prophecy. Its discipline, forcing the estimate to stand next to the distribution of what similar work actually cost, is the strongest single counterweight to the inside view and its companion pressures. The software industry has its own reference-class evidence, with caveats attached: the Standish Group’s CHAOS research, widely cited in the 2000s, reported in its 2004 study that IT projects overran their budgets by 43 percent on average, that 71 percent of projects came in over budget, over time, or over scope, and that waste in the United States alone ran to roughly 55 billion dollars a year. The CHAOS methodology is industry research with contested definitions, and it should be read as directional, not authoritative. The point is the habit: the honest estimator goes looking for the distribution instead of trusting the simulation.

At BlueLine the outside view has a home: the consortium’s reference database of comparable corridors, plus the city’s own portfolio, per lane-kilometre and per station, with the differences stated. The estimate for the eastern segment is not “what our engineers think the flood plain will cost” but “what flood-plain segments in the class have cost, adjusted for the differences our engineers can price, plus the uncertainty they cannot.” The reference class is built before the estimate, not after, and it is the first check the estimate must survive.

What makes an estimate credible

The credibility machinery is three small documents, and any of them alone is better than a single point.

The basis of estimate is the minimum viable tool: one page that says what was estimated, the scope definition and its version, which method produced the number and why, what data fed it and from when, what assumptions it carried, what it excluded, the range and its confidence, the maturity of the estimate, the owner, and the review date. The basis is what makes the number defeatable: a reader can check the method, question the data, and price the assumptions.

The assumptions log is the estimate’s nervous system: one row per assumption, the assumption, why it matters, who owns it, when it will be tested, and what happens to the number if it fails. The log exists because the estimate is a bundle of assumptions, and the bundle is the real subject of every estimate review. When an assumption changes, the estimate changes, and the change is a decision, not a surprise: the flood-plain assumption has an owner and a trigger date, and when the trigger fires, the estimate moves by the priced amount, through the change control of chapter 39, and the forecast updates.

The confidence statement adds the humility the range cannot: the answer to “what would have to be true for the final number to land 10 percent lower, or 10 percent higher?” The range says how wide the number is; the statement says what would have to happen at either edge, so the estimate can be examined instead of defended.

The estimate that learns

The last discipline is the cadence: an estimate is a living instrument, and its updates are the project’s learning.

The re-estimation triggers are the tailoring triggers of chapter 13, wearing numbers. A gate arrives, and the estimate is rebuilt for the next stage with the evidence the gate produced. An assumption breaks — the drainage approval slips, the integration spike lands — and the affected part of the estimate is re-priced immediately. Retired uncertainty arrives, the prototype proves the method, the class-three estimate becomes a class-two estimate, and the cone narrows by decision, not by decree. The forecast updates on the cadence, every month at BlueLine, every Sprint at KijaniPay; the commitment changes only through governance, with the full visibility of chapter 8. The distinction is the four-words discipline applied over time: the forecast moves with evidence, the commitment moves with authority, and the two never merge.

The failure is the frozen cone: the estimate written in month three, defended in month seventeen, and never re-examined, while the evidence accumulated around it. The frozen cone is not stability; it is the project pretending that knowledge did not grow. The repair is the re-estimation gate, scheduled like the other gates, where the estimate is rebuilt from current evidence and the difference from the old number is presented as a finding: here is what we learned, here is what it cost, here is the new range. Re-estimation is not failure, and the project that treats it as failure will stop learning at the exact moment learning is cheapest, which is before the surprise.

At BlueLine the re-estimation is already scheduled: the second-half authorization is the gate, and the integrated estimate going into it must price the seam that chapter 14 found, the integration and rehearsal work that lived in the capability view and in no contract. The estimate that leaves it out is the disappearing 100 percent in financial form, precise, complete-looking, and missing the work that will actually decide the opening date.

The six ways a number lies

The failure patterns of estimation deserve to be named as characters, because each one is produced by competent people doing what the room rewarded.

The precise liar is the estimate with more significant digits than knowledge: the 1,486.7 million units, the 2,347.3 hours, the model output quoted to the decimal while its inputs were ranges. The tell is the decimal where the range should be. The cost is the manufactured certainty that chapter 3 called the most expensive error, dressed in spreadsheet clothes. The repair is the range and the basis: every number reports how wide it is and where it came from.

The sandbag is the hidden contingency: the honest range that could not be said aloud, added to the point in secret. The tell is the estimate that does not move, the project that is consistently slightly under budget, the analyst who says “I built in some slack” and cannot say how much. The repair is the declared pocket: the range is said aloud, the contingency is labeled, and the sandbag is converted into the two-pocket structure it was always pretending to be.

The anchored room is the meeting where the first number wins: the sponsor’s target spoken first, every later number an adjustment from the anchor. The tell is the estimating session where the discussion is about the anchor, the last project, the budget line, the boss’s guess, and never about the work. The repair is the simultaneous reveal: estimates are written and shown together, before the discussion, in the planning-poker pattern, so that the first number is not the first word.

The target in an estimate’s coat is the number that must be hit, dressed as analysis: the date the mayor promised, the budget the council appropriated, presented as the estimate, with the work bent to fit it. The tell is the phrase “we will make it work” where the estimate should be. The cost is the death march, the project that walks itself to exhaustion keeping a promise no analysis supported. The repair is the four-words discipline: the target is named as a target, the estimate as an estimate, and the gap between them is made into a decision about scope, time, cost, and risk, the decision of chapter 2’s trade-offs, instead of a silent fiction.

The frozen cone is the estimate that never learns: written once, defended forever, re-examined never. The tell is the number that has not moved in months while the project moved around it. The repair is the re-estimation gate.

The orphan is the estimate with no basis, no owner, no assumptions: the number in the deck, the figure in the email, the line in the plan with no address. The tell is the question “where did that number come from?” followed by a pause. The repair is the one-page basis, attached to the number, making it defeatable, checkable, and owned.

The six characters share one root: each one replaced the decision the estimate was supposed to serve with a performance. The estimate exists to make a better decision possible, and the honest range, the declared pockets, the named assumptions, and the scheduled re-estimation are the practices that keep the number in service of the decision rather than in service of the room.

The assistant’s narrow lane

The assistant has a genuine and bounded role in estimating, and its boundary is the book’s standing boundary. It can draft the first pass of a basis of estimate from provided scope documents, contracts, and prior estimates, labeling what came from where. It can check the arithmetic of a three-point estimate, verify the PERT formulas, reconcile the sums against the work package dictionary of chapter 14, and flag totals that do not close. It can generate scenario ranges and sensitivity tables from the assumptions log, showing how the estimate moves when each assumption fails. It can search provided historical data for reference-class candidates and summarize the distribution of comparable outcomes. The source data is the approved, non-confidential project context, redacted before prompting. The draft stays a draft until a named owner verifies it, and the audit record says what was generated, from what, checked by whom, and decided by whom.

Three boundaries are worth naming, because estimation is where the machine’s fluency is most seductive. First, the scope boundary is not discoverable from the documents: what sits inside the estimate and what belongs to operations or to another program is a decision argued with the sponsors and the operational owners, and the machine’s proposed boundary is a hypothesis until the people confirm it. Second, the data is not self-grounding: a fluent model can produce a confident number with no provenance, and an estimate without provenance is the precise liar generated at speed. Every input must carry its source and date, and every reference-class candidate must be verifiable. Third, the commitment is not delegable: the machine can draft the range, and the accountable person makes the commitment, because a commitment is an acceptance of obligation under the governance of chapter 8, and no draft is an obligation. The verification is the room: the leader tests the draft basis against the scope, the contracts, the assumptions log, and the reference class, the way this chapter’s opening room should have. The machine accelerates the drafting; the boundary, the provenance, and the commitment stay human.

Practice

One. A quick classification. Label each statement as estimate, target, commitment, or forecast, and say which owner each one implies. (a) “The corridor opens in month twenty-four; the mayor promised the council.” (b) “Our current projection, from the integrated estimate, is month twenty-five, with an 80-percent range of month twenty-four to twenty-seven.” (c) “Lena accepted a completion obligation of month twenty-five at the second-half authorization.” (d) “We want the remaining forward cost under 1,450 million units, and we will design the scope to it.” (e) “As of this month’s steering committee, our opening forecast is month twenty-five, and the 80-percent range has narrowed to month twenty-four and a half to twenty-six.”

(a) is a target: the mayor’s ambition, owned by the sponsor, and the room must not confuse it with analysis. (b) is a forecast: the current evidence-based expectation, owned by the analysts, updated as evidence arrives; the 80-percent range is its honest shape. (c) is a commitment: an accepted obligation, owned by the accountable person, changeable only through governance. (d) is a target wearing planning clothes: the ceiling the organization wants, and the sentence “we will design the scope to it” is exactly the target in an estimate’s coat unless it is named as a target and the scope consequence is priced. (e) is the forecast again, a month later, and the narrowing range from (b) to (e) is learning, not failure: the cone narrowed because the evidence retired uncertainty, and the forecast moved with it. The common error is treating (a) as an estimate: it turns the mayor’s ambition into the analyst’s number, and the analyst’s number into a commitment nobody authorized.

Two. A field drill: build the range. Take one bounded deliverable you know, one work package from your own chapter 14 structure, and produce the honest number for it. (a) Write the optimistic, most likely, and pessimistic durations with the pessimistic defined as a genuine tail, not a bad-but-quotable day. (b) Compute the PERT expected value and the rough standard deviation, and write the one-standard-deviation range. (c) Write the confidence statement: the range, the three largest assumptions, who owns each, and what would have to be true to land 10 percent early or 10 percent late. (d) Then run the meeting you actually sit in: who spoke the first number, and how far did the final estimate land from the anchor?

The drill succeeds when the arithmetic is checkable, the pessimistic case is uncomfortable, and the assumptions are owned. The most common failure is the trimmed pessimistic: the tail set at the point the room can tolerate, not the point the evidence supports, which shrinks the range to theater. The second is the unowned assumption: the swing factor that moves the number most, named with no owner, which means it will move the number without a decision. The anchor audit is the part people skip, and it is the part that transfers: the first number in your room is doing more work than the evidence.

Three. A decision room: the number the city asked for. It is month seventeen at BlueLine, and the authorization pack is due in five days. The mayor’s office has repeated the demand: one figure for the council, no ranges. The choices on the table: present the 1,390 million-unit base as the figure, with the range in a technical annex; present the 80-percent range of 1,280 to 1,500 as the headline; or present the two-pocket structure, 1,390 base, 110 million-unit contingency, 1,570 ceiling, as the authorization ask, and force the council to appropriate the ceiling. Decide what Lena should take to the mayor’s office, how she should handle the accusation that a range is indecision, and what she should refuse to do. Defend the trade-off.

The defensible answer presents the range as the truth and the two-pocket structure as the authorization: 1,390 base, 110 contingency managed by the project against the named risks, and the band to 1,570 as the city’s reserve, released only by re-authorization. The range is not indecision; it is the difference between a number and a bet. The political framing shows the council what it is buying: a corridor with an 80-percent chance of landing inside the stated band, plus the mechanism that absorbs the tail. The refusal is the precise liar: Lena does not submit a single point as the official figure, because it is a fiction that will be rediscovered as an overrun, with the council’s trust and the project’s credibility in the same envelope. Reasonable but risky: presenting 1,390 as the figure with the range in an annex — defensible only if the briefing is explicit that 1,390 is the base, that the range is the estimate, and that the annex will be read. That is a judgment about the council’s reading habits, not about the arithmetic. The unsafe choices: the precise single point, the whole failure pattern of this chapter in one page; and the range with no pocket structure, which tells the truth about width and nothing about who absorbs the tail, leaving the council to imagine its own contingency — which it will set too low.

Four. The mastery drill: the precise estimate and the wide one. Two estimators present the same work. Estimator A says: “The integration takes 9.3 weeks, 95 percent sure.” Estimator B says: “The integration takes 9 to 17 weeks, most likely 10 to 12, and the range is driven by the legacy interface, which we have not scoped.” Explain why B’s estimate may be more credible than A’s, what A’s number is actually doing, and how each estimate should be treated in the decision. Then say what evidence would let B narrow to A’s width honestly.

B’s estimate is more credible because it states the uncertainty and its source: the range is the estimate, the width is information, and the assumptions are visible. A’s number is a point with a probability attached, and the probability is the problem: a 95-percent claim requires a distribution, and a distribution requires evidence A has not shown. The claim is either borrowed from a model A cannot defend or invented to satisfy the room. A is doing persuasion, not estimation: it reports confidence as a property of the speaker rather than of the evidence — the precise liar with a probability attached. In the decision, B’s estimate supports the two-pocket structure: the contingency is sized from the range, and the trigger is the interface scoping. A’s estimate supports only the illusion that the integration is under control, which the risk register of chapter 22 will have to correct later. The evidence that lets B narrow honestly is the evidence the cone requires: scope the legacy interface, run the integration spike, retire the uncertainty by decision, and re-estimate from what the spike proved. The unsafe response is choosing A because the project needs a number: the project always needs a number, and the honest number is the range. The wrong reading is treating B’s estimate as a failure of estimating; it is the opposite, and the distinction is the whole chapter.

Five. The transfer question. Look at the estimate your project is currently carrying. Where is its basis, and does it name its method, its data, its assumptions, and its exclusions? Which of the six characters does it resemble: the precise liar, the sandbag, the anchored room, the target in an estimate’s coat, the frozen cone, or the orphan? And if you added one page this week, the basis of estimate or the assumptions log, which would change the next decision more?

The durable principle: an estimate is a decision instrument, and its honesty is the width it reports, the assumptions it names, the pockets it declares, and the cadence at which it learns. The precise single number is not an estimate; it is a bet, and the bet’s odds are hidden. The most common next failure is quieter than the six characters named here: the project produces an honest range at the gate, the gate passes, and then the range is forgotten, the pockets merge, the assumptions log is archived, the re-estimation gate becomes a formality, and the project is once again managing a single point that no one believes and everyone defends. The range must be carried into the plan itself, the integrated roadmap and management plan of the next chapter, which is where the estimate becomes a commitment with a shape: the baselines, the milestones, the contingency, and the learning cadence all built on the honest number, so that the honesty survives contact with the plan.

Notes

  • BlueLine Urban Mobility Program is a composite case created for this book; no real city, corridor, contractor, or people are depicted. The month-seventeen authorization scene, Marta Reyes, the analyst’s precise figure of 1,486.7 million units, and the forward-base arithmetic are author-created illustrative material consistent with the earlier chapters: the 24-month program, the 2,400 million-unit capital envelope and the 140 million-unit mitigation line from chapter 7, the roughly 1,000 million units already committed by month ten, and the integrated planning session at month sixteen from chapter 14. The forward arithmetic here is 2,400 minus 1,150 committed, 1,250 remaining as engineered, plus 140 for mitigations, a forward base of 1,390; the 80-percent range of 1,280 to 1,500, the 110 million-unit contingency, and the 1,570 ceiling are author-created teaching numbers from the corridor’s fictional risk model, not an audit, and any reader can reproduce them from the text. The three-point ticketing example (5, 9, 17 weeks) and the expected-value vendor example are author-created teaching numbers.
  • The cone of uncertainty follows Barry W. Boehm, Software Engineering Economics (Englewood Cliffs, NJ: Prentice Hall, 1981), where Boehm presents the “funnel curve” and places the early uncertainty at a factor of about four in either direction (p. 311); the name “cone of uncertainty” was first used for the concept by Steve McConnell in Software Project Survival Guide (Redmond, WA: Microsoft Press, 1997), and the practical treatment follows McConnell, Software Estimation: Demystifying the Black Art (Redmond, WA: Microsoft Press, 2006). The chemical-industry precursor follows H. Carl Bauman, “Accuracy Considerations for Capital Cost Estimation,” Industrial and Engineering Chemistry 50(4) (1958), with the AACE Bulletin “Estimate Types” (November 1958); the American Association of Cost Engineers is now AACE International. Chapter 13 of this book carries the same cone at life-cycle level; the two chapters use it for different decisions.
  • The planning fallacy follows Daniel Kahneman and Amos Tversky, “Intuitive Prediction: Biases and Corrective Procedures,” in Judgment Under Uncertainty: Heuristics and Biases (Cambridge: Cambridge University Press, 1982; original technical report 1977; also published 1979 in TIMS Studies in the Management Sciences), and the expanded definition follows Dan Lovallo and Daniel Kahneman, “Delusions of Success: How Optimism Undermines Executives’ Decisions,” Harvard Business Review 81(7), July 2003. The thesis study follows Roger Buehler, Dale Griffin, and Michael Ross, “Exploring the ‘Planning Fallacy’: Why People Underestimate Their Task Completion Times,” Journal of Personality and Social Psychology 67(3), 1994, which reported average predicted completion of 33.9 days against actual 55.5 days.
  • Anchoring follows Amos Tversky and Daniel Kahneman, “Judgment under Uncertainty: Heuristics and Biases,” Science 185(4157), 1974, including the roulette-wheel demonstration; the Gandhi-age demonstration follows Fritz Strack and Thomas Mussweiler, “Explaining the Enigmatic Anchoring Effect,” Journal of Personality and Social Psychology 73(3), 1997.
  • Strategic misrepresentation and the infrastructure evidence follow Bent Flyvbjerg, Mette K. Skamris Holm, and Søren L. Buhl, “Underestimating Costs in Public Works Projects: Error or Lie?” Journal of the American Planning Association 68(3), 2002, based on 258 transportation infrastructure projects worth about US$90 billion: costs underestimated in almost 9 of 10 projects, an 86 percent likelihood of actual cost exceeding estimate, actual costs on average 28 percent higher, and average escalation of 44.7 percent for rail, 33.8 percent for fixed links, and 20.4 percent for roads. The quoted phrase “that is, lying” is the authors’ own summary of strategic misrepresentation in the paper’s abstract; the term itself has an earlier history in public budgeting research (Larry R. Jones and Kenneth J. Euske, “Strategic Misrepresentation in Budgeting,” Journal of Public Administration Research and Theory 1(4), 1991).
  • The software-industry reference evidence follows the Standish Group, CHAOS Report (West Yarmouth, MA, 2004), which reported a 43 percent average cost overrun, 71 percent of projects over budget, over time, or over scope, and about US$55 billion in annual waste; the CHAOS methodology and definitions are contested in the research literature, and the figure is presented here as directional industry evidence with disclosed limits, not as a certified statistic.
  • PERT follows the primary historical record: the Program Evaluation and Review Technique was developed for the United States Navy’s Polaris missile program by the Navy’s Special Projects Office, Lockheed Aircraft, and Booz Allen Hamilton, and made public in 1958; the expected-duration formula (optimistic + 4 times most likely + pessimistic) divided by 6, with the rough standard deviation (pessimistic minus optimistic) divided by 6, is the technique’s standard convention. The expected-value interpretation in this chapter (two-thirds of outcomes within one standard deviation under a normal approximation, and about a quarter chance of exceeding a value two-thirds of a standard deviation above the mean) is the author’s worked interpretation for teaching, not a claim about the exact distribution.
  • The Delphi method follows Norman Dalkey and Olaf Helmer, “An Experimental Application of the Delphi Method to the Use of Experts,” Management Science 9(3), 1963, for the RAND Corporation work; wideband Delphi follows Barry Boehm and John A. Farquhar’s development in the 1970s, popularized in Boehm, Software Engineering Economics (1981); planning poker follows James Grenning, “Planning Poker” (2002), popularized by Mike Cohn, Agile Estimating and Planning (Upper Saddle River, NJ: Prentice Hall, 2005).
  • Little’s Law follows John D. C. Little, “A Proof for the Queuing Formula: L = λW,” Operations Research 9(3), 1961; the practitioner treatment of throughput-based forecasting follows Daniel S. Vacanti, Actionable Agile Metrics for Predictability: An Introduction (Neptune Township, NJ: DZone, 2015). Chapter 17 of this book builds the full flow machinery with the law’s assumptions stated.
  • Reference class forecasting follows Lovallo and Kahneman, “Delusions of Success” (2003), and the practical guidance Bent Flyvbjerg and COWI developed for the UK Department for Transport, Procedures for Dealing with Optimism Bias in Transport Planning (June 2004). The Edinburgh Tram Line 2 numbers follow the documented account of the October 2004 review by Ove Arup and Partners Scotland: forecast total capital cost 320 million pounds with 25 percent contingency; reference-class 50th percentile 357 million pounds (40 percent contingency) and 80th percentile 400 million pounds (57 percent contingency); reviewers’ warning that the forecasts were likely too low; final outturn about 776 million pounds (about 628 million in 2004 prices), opened May 2014, three years late. The numbers as reported in the literature are reproduced here with their sources; the chapter’s account is a summary of a documented case, not an endorsement of any party’s conduct.
  • The Scrum vocabulary follows the current Scrum Guide, November 2020 version, by Ken Schwaber and Jeff Sutherland: each artifact carries a commitment (Product Goal for the Product Backlog, Sprint Goal for the Sprint Backlog, Definition of Done for the Increment), and the work selected for a Sprint is discussed as a forecast, with confidence grounded in past performance, upcoming capacity, and the Definition of Done.
  • The estimate classification follows AACE International’s recommended practice 18R-97, “Cost Estimate Classification System,” which ties estimate classes (class five for concept work through class one for near-definitive engineering) to the maturity of project definition; readers should consult the current edition of the recommended practice for the accuracy ranges, which vary by industry and by version.
  • ISO 21502:2020, “Project management: Guidance on project management,” addresses planning and estimating at a general level without prescribing a single estimating technique; the four-word vocabulary, the two-pocket structure, the basis of estimate, and the confidence statement are the author’s method-neutral working instruments, and this book is independent of ISO, PMI, AACE, Scrum, and all named framework owners. The composite cases, including KijaniPay, remain author-created illustrative material consistent with the facts established in chapters 6, 12, and 14.