All articles
6 min readentrepreneurship / productivity

Planned 33.9 days, took 55.5: estimate from past projects

Most scoping calls I run end on the same question: how many weeks. I answer with a story about how the work will go, and the story I tell is the shortest plausible one. Roger Buehler put a number on that habit in 1994 by asking psychology students when they would finish their thesis. They said 33.9 days. They took 55.5 1.

What the research shows.

Buehler, Dale Griffin and Michael Ross followed 37 students through their senior thesis. Mean prediction: 33.9 days. Mean actual: 55.5. About 30% finished inside their own forecast. The same students also gave a worst case, the date they would hit "if everything went as poorly as it possibly could": 48.6 days. The average student overshot the pessimistic number by a week. Think-aloud protocols in the same paper explain the mechanism. People narrate a plan for the task in front of them and leave their own track record out of the narration. The optimism is specific to the self, since those students predicted other people's completion times without it 1.

The pattern scales up. Kjetil Moløkken and Magne Jørgensen reviewed the surveys on software effort estimation and found 60% to 80% of projects overrunning effort, schedule or both, with effort overruns averaging 30% to 40% 2. Bent Flyvbjerg's team measured 258 transport infrastructure projects across 20 nations, worth around US$90 billion. Actual costs landed 28% above estimate on average, and the gap varied by type: 44.7% for rail, 33.8% for bridges and tunnels, 20.4% for roads 3. His database of more than 16,000 projects puts 47.9% on budget, 8.5% on budget and on time, and 0.5% on budget, on time and delivering the promised benefits 4.

Four corrections have been tested.

Force the past into the forecast. In the fourth study of the 1994 paper, students first named the date they would finish if they finished as far ahead of the deadline as they typically do, then wrote a plausible scenario, drawn from their own history, that produced that date. The optimistic bias disappeared 1.

Unpack the task. Justin Kruger and Matt Evans had people break down holiday shopping, getting ready for a date, formatting a document and preparing food into components before predicting. Listing the parts cut the planning fallacy by more than half 5.

Plan backward. Jessica Wiese, Buehler and Griffin compared forward planning, first step to last, with backward planning that starts at the finished result and walks the steps in reverse. Backward planners predicted longer completion times across four studies, and in the fourth those longer predictions were the less biased ones 6.

Forecast from other people's projects. Flyvbjerg's reference class forecasting drops the narrative estimate and reads the distribution of outcomes from comparable finished projects. The UK Department for Transport published the method in August 2004 and required local authorities seeking funding to apply optimism bias uplifts. Edinburgh Tram Line 2 was the first live case: the promoters carried a base cost of £255 million plus 25% for contingency, while a reference class of 46 comparable rail projects put the 50th percentile at £357 million, a 40% uplift, and the 80th percentile at £400 million, a 57% uplift 7.

The bias reverses on short tasks.

Torleif Halkjelsvik and Jørgensen reviewed the judgment-based time-prediction literature in Psychological Bulletin. Underestimation dominates the studies coming out of engineering and management. It does not dominate the psychology experiments, which mostly run on tasks of minutes rather than months 8. Padding a 20-minute chore by 40% costs you an afternoon. Uplifts belong on work measured in weeks, with handovers and other people in the chain.

Edinburgh marks the other limit. The line opened three years late with an outturn cost of £776 million, £628 million in 2004 prices, above even the 80th percentile forecast. Flyvbjerg's report for the tram inquiry notes that the guidelines called for uplifts at the decision to build, a stage the project had not reached when that forecast was made 9. A reference class narrows the error. It does not close it.

One more trap sits inside the unpacking advice. Halkjelsvik's experiments on the summation fallacy show that adding subtask estimates carries the bias of each item into the total: sum optimistic estimates and the total comes out low, sum conservative ones and it comes out high 10. Treat the sum as an input, not the answer.

The protocol.

  1. Write your gut estimate and date it before you look at anything else. It becomes the number you audit later.
  2. List the last 5 comparable projects your team finished, with their actual durations. Same team size, same kind of client, same definition of done. Take the median of the actuals as your starting forecast.
  3. With fewer than 5 to draw on, borrow a published multiplier for the domain: around 1.3 on software effort 2, around 1.2 on build-style work 3.
  4. Unpack the work into components, review cycles, handovers and waiting time included 5. Estimate each at its typical duration, not its best case, then treat the sum as a floor rather than the answer 10.
  5. Re-plan backward from the delivery date and keep the longer of the two numbers 6.
  6. Publish two dates: a 50% date and an 80% date, the way the uplift tables do 7. Commit externally to the 80% one and work toward the 50%.
  7. Ask a colleague who is not on the project for their estimate, and record it. People forecast other people's timelines with less optimism 1.
  8. Log three fields per project: estimated, promised, actual. At the fifth row you stop borrowing multipliers and read your own median.

Start the log with the project you are in the middle of right now: write the date you predicted, the date you promised, and the date you currently expect.

Sources.

  1. Buehler, R., Griffin, D., & Ross, M. (1994). Exploring the "planning fallacy": Why people underestimate their task completion times. Journal of Personality and Social Psychology, 67(3), 366-381. doi.org/10.1037/0022-3514.67.3.366
  2. Moløkken, K., & Jørgensen, M. (2003). A review of surveys on software effort estimation. International Symposium on Empirical Software Engineering (ISESE 2003). doi.org/10.1109/ISESE.2003.1237981
  3. Flyvbjerg, B., Skamris Holm, M. K., & Buhl, S. L. (2002). Underestimating costs in public works projects: Error or lie? Journal of the American Planning Association, 68(3), 279-295. doi.org/10.1080/01944360208976273
  4. Flyvbjerg, B., & Gardner, D. (2023). How Big Things Get Done. Currency. flyvbjerg.plan.aau.dk
  5. Kruger, J., & Evans, M. (2004). If you don't want to be late, enumerate: Unpacking reduces the planning fallacy. Journal of Experimental Social Psychology, 40(5), 586-598. doi.org/10.1016/j.jesp.2003.11.001
  6. Wiese, J., Buehler, R., & Griffin, D. (2016). Backward planning: Effects of planning direction on predictions of task completion time. Judgment and Decision Making, 11(2), 147-167. journal.sjdm.org
  7. Flyvbjerg, B. (2008). Curbing optimism bias and strategic misrepresentation in planning: Reference class forecasting in practice. European Planning Studies, 16(1), 3-21. doi.org/10.1080/09654310701747936
  8. Halkjelsvik, T., & Jørgensen, M. (2012). From origami to software development: A review of studies on judgment-based predictions of performance time. Psychological Bulletin, 138(2), 238-271. doi.org/10.1037/a0025996
  9. Flyvbjerg, B. (2018). Report for the Edinburgh Tram Inquiry. University of Oxford. arxiv.org/abs/1805.12106
  10. Halkjelsvik, T. (2022). When 2 + 2 should be 5: The summation fallacy in time prediction. Journal of Behavioral Decision Making, 35(3), e2265. doi.org/10.1002/bdm.2265