Estimation techniques & sizing
No estimation technique removes uncertainty. The technique should match the evidence available and the decision being made. This chapter is ordered from the methods you will actually reach for most weeks down to the specialist ones you reach for rarely — that order is itself the lesson.
Evidence available? → historical data: analogy/calibration | multiple experts: Delphi | uncertain range: three-point | relative backlog: relative sizing → record assumptions → compare with actualsExpert judgmentCommon
Expert judgment uses relevant hands-on experience to size work directly: an engineer who tested a similar integration before estimates the new one from that experience, adjusted for known differences.
Practical use: fastest available method; state the comparable experience out loud ("this is like the SSO integration, but with two more identity providers") so the reasoning can be challenged, not just the number.
Caveat: a single expert's judgment carries that person's blind spots. For anything with real budget consequences, pair it with a second opinion or a documented reference case.
Bottom-up estimationCommon
Decompose the work into smaller testing activities, estimate each one, then sum.
Checkout testing
Requirements review 4 h
Test design 8 h
Test data 4 h
API testing 8 h
UI testing 10 h
Payment integration 10 h
DB validation 3 h
Mobile verification 4 h
Regression 8 h
Reporting 2 h
------------------------------
Base effort 61 hPractical use: reach for this when scope is reasonably understood, work can genuinely be decomposed, and a QA Lead needs to expose hidden work before committing to a number.
Caveat: a detailed table is not automatically accurate. If the decomposition misses environment work, defect re-testing, test-data preparation, external dependencies, review or reporting, the sum is still wrong — it just looks precise.
Planning Poker, T-shirt sizing and relative sizingCommon
Relative sizing compares items with one another rather than converting each directly to time. Planning Poker is one collaborative way to expose different assumptions; T-shirt sizes offer coarser categories when precision is not justified.
Practical use: use a stable reference item and discuss outliers. Keep sizing separate from staffing and calendar scheduling until capacity and historical delivery evidence are considered.
Caveat: story points or T-shirt sizes are local scales. Converting them with an industry-wide hours-per-point constant defeats their purpose.
What Scrum actually requires—and does not requireCommon
The current official Scrum Guide defines Scrum without prescribing story points, Planning Poker or velocity as mandatory practices. Teams may use useful sizing and forecasting techniques, but they should not be presented as requirements of Scrum.
Practical use: when a process rule is defended with 'Scrum requires it', verify the current guide. Keep local working agreements clearly identified as local.
Caveat: the absence of a mandated technique does not mean estimation is useless; it means the team can choose an evidence-based method appropriate to its context. A team that derives a local story-point-to-hour ratio for its own forecasting is doing empirical calibration, not redefining what a story point means.
Optimistic, most-likely and pessimistic estimatesCommon
Three-point estimation describes uncertainty with optimistic (O), most-likely (M) and pessimistic (P) cases. The value is primarily in forcing assumptions about favourable, normal and adverse conditions into the discussion — the arithmetic in the next section is secondary to this.
Practical use: for each point, state the scenario: what must be true for O, what you expect for M, and which plausible risks drive P.
Caveat: O and P should not be fantasy extremes. They should represent plausible bounds for the decision context.
Top-down estimationLess common
Estimate the overall effort or budget first, then allocate it across activities or components.
scenario example — not a universal standard
Available QA budget = 100 h
Functional testing 35 h
Integration 20 h
Regression 25 h
Automation 15 h
Reporting 5 hPractical use: useful for early portfolio planning, rough budgeting, and very early release planning when details are not yet available.
Caveat: top-down estimation can hide missing work. It is weakest when used backwards — "we have 5 days, therefore testing takes 5 days" — because capacity is a constraint, not evidence of the effort actually required.
Parametric estimationLess common
Estimate work from a calibrated unit rate.
Historical execution rate: 5 min per comparable test
400 tests × 5 min = 2,000 min ≈ 33.3 h
Historical API-test design rate: 20 min per comparable case
120 cases × 20 min = 40 hPractical use: strongest when work is repetitive, units are genuinely comparable, the team has historical actuals, and the parameter gets recalibrated periodically.
Caveat: a rate copied from another company or another class of work is not a parameter — it is an assumption wearing a parameter's clothes. Calibrate locally before trusting it.
A worked PERT-style calculationLess common
A common PERT-style weighted estimate uses E = (O + 4M + P) / 6. This is a model, not a law: its usefulness depends on the quality of O, M and P and on whether the underlying assumptions fit the work.
Practical use: scenario example: O=4 days, M=6 days, P=11 days gives E=(4+24+11)/6=6.5 days. Keep the original range and assumptions beside the weighted value; do not report 6.5 as certainty.
Caveat: a weighted point estimate compresses information. If the decision depends on tail risk, retain the range or use a probabilistic forecast rather than hiding it.
scenario example — not a universal standard
formula: E = (O + 4M + P) / 6
numerator: O + 4M + P (weighted scenario estimates)
denominator: 6 (the PERT-style weighting constant)
scope: one defined work item under the stated assumptions
window: estimate is current until scope/evidence changes
target: not applicable
owner: estimating team / decision owner
decision: use the weighted value with O–P range as planning evidence, then recalibrate from actuals
example: O=4 days, M=6 days, P=11 days → E=6.5 daysWhen the PERT formula gives false comfortLess common
PERT-style arithmetic can look scientific even when the inputs are guesses, dependencies are correlated or the distribution is asymmetric. Precision in the formula cannot compensate for weak evidence.
Practical use: use the method as a structured conversation and compare estimates with actuals. Escalate to scenario or Monte Carlo forecasting when multiple uncertain items interact.
Caveat: reporting two decimal places from coarse O/M/P inputs is false precision. Learning three-point thinking matters more than memorizing the arithmetic.
Wideband DelphiLess common
Several experts estimate independently, then discuss the spread and re-estimate until it converges or the disagreement itself becomes informative.
Practical use: clarify the same scope for everyone, collect estimates privately, reveal the range, ask the high and low estimators to explain their assumptions, update scope/evidence, then estimate again.
Caveat: consensus is not the goal if disagreement reflects real uncertainty. Preserve the reasons for the spread rather than forcing everyone to one number.
Function points and use-case pointsRare
Function points are a standardized functional-size family; ISO/IEC 20926 specifies the IFPUG method. Use-case points are a different estimation approach built around use cases and adjustment factors and should not be presented as the same standard.
Practical use: use formal sizing only where the organization has a reason to maintain the counting rules and enough historical delivery data to calibrate size to effort or duration.
Caveat: functional size is not effort. A function-point count becomes an effort forecast only through context-specific productivity evidence.
Test points as a calibrated sizing modelRare
A 'test points' model is best treated as a locally defined sizing heuristic that weights testing drivers such as complexity, interfaces, data, quality risk or execution conditions. There is no single universal conversion to hours.
Practical use: define the factors, weights and counting rules explicitly, then calibrate the resulting score against completed local work. Version the model when rules change.
Caveat: a weighted spreadsheet is not validated merely because it produces a number. Without calibration and error tracking, the weights are opinions disguised as a model.
Choosing a sizing model: calibration before conversionCommon
Choose sizing methods by decision need, data availability, repeatability and calibration cost. Coarse relative sizing can be enough for prioritization; formal functional sizing may suit organizations that need repeatable cross-release measurement.
Practical use: run a candidate model on historical items before using it for commitments. Measure its forecast error and compare it with simpler baselines.
Caveat: a more complicated model is worse if it is not more accurate, understandable or decision-useful than a simple reference-class estimate.
Summary
- This chapter covers 13 required concepts while keeping tool/formula details tied to a practical decision.
- Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
- Claims that depend on a standard or product are grounded in the source registry below.