GimmeJob
Sign in
QA metrics & estimation · Chapter 05 / 08

QA Metrics & Estimation

Estimation foundations & decomposition

Good estimation is not the production of a precise-looking number. It is a decision-support process that makes scope, assumptions, dependencies and uncertainty visible. The most useful first move is usually decomposition, not arithmetic.

Request → clarify scope → decompose work → expose dependencies/uncertainty → choose evidence/method → range + assumptions → re-estimate when evidence changes

Estimate vs commitmentCommon

An estimate describes what is currently believed about effort, duration or completion under stated assumptions. A commitment is a decision to pursue an outcome under constraints. Treating the estimate as a promise destroys information about uncertainty.

Practical use: State the estimate, confidence/range, assumptions and decision date separately. A stakeholder can still make a commitment, but the evidence remains visible.

Caveat: Pressure to produce one date is a governance problem, not evidence that uncertainty disappeared.

Effort vs durationCommon

Effort is the amount of work, commonly expressed in person-hours or person-days. Duration is elapsed calendar time. Dependencies, waiting, environments, review queues and parallelism mean the two are not interchangeable.

Practical use: Estimate effort by work package, then build duration from capacity, sequence, dependencies and waiting. Show both when planning test work.

Caveat: Dividing 20 person-days by four people does not automatically produce five calendar days; some work cannot be parallelized and adding people has coordination cost.

Point estimate vs range and confidenceCommon

A point estimate hides uncertainty in a single value. A range communicates plausible outcomes, while confidence explains how strongly the available evidence supports that range.

Practical use: Prefer a range when uncertainty is material: for example, '6–9 days given stable test data and one environment'. Add which assumptions dominate the width.

Caveat: Do not attach a percentage confidence unless the team has a defensible method for producing it. Verbal confidence can be clearer than invented statistics.

Contingency vs hidden paddingCommon

Contingency is an explicit allowance for identified uncertainty or risk. Padding is concealed extra time embedded in an estimate. Explicit contingency can be reviewed and retired; hidden padding cannot.

Practical use: Tie contingency to named risks such as unstable environments, migration uncertainty or external certification. Revisit it as those risks change.

Caveat: A blanket '20% buffer' without evidence is easy to normalize into the baseline and stops communicating why uncertainty exists.

Capacity constraints and knowing when to re-estimateCommon

Capacity constrains what can be completed during a period; it does not change the intrinsic effort of a task. Re-estimation is warranted when scope, assumptions, evidence, dependencies or available capacity materially change.

Practical use: Define triggers before work starts: scope change, blocked environment beyond a threshold, major defect discovery, dependency slip, or meaningful difference between actuals and assumptions.

Caveat: Re-estimating every day without new evidence creates churn; refusing to re-estimate after the model changed creates false certainty.

Work breakdown structure and hidden testing workCommon

A work breakdown structure decomposes an outcome into manageable work packages. For testing, hidden work often includes analysis, test data, environment setup, automation support, defect investigation, re-testing, reporting and coordination.

Practical use: Build the WBS around deliverable work, then check every package for prerequisites and completion evidence. Compare it with similar completed work.

Caveat: A WBS is not useful if it only restates feature names. Decompose until the work and uncertainty become estimable.

QA decomposition patterns: functional, component, test-type, platform, risk-basedCommon

WBS answers how far down to break work; these five patterns answer along which axis. The same feature usually needs more than one at once.

Functional — use when product behavior is the main planning boundary:

Checkout
├── Address
├── Delivery
├── Promo code
├── Tax
├── Payment
└── Confirmation

Component — use when architecture and integrations drive test risk:

Checkout
├── Web UI
├── Backend API
├── Payment service
├── Database
├── Notification service
└── External payment provider

Test-type — use when specialized work must be planned or assigned:

Checkout
├── Functional
├── API
├── Integration
├── Exploratory
├── Regression
├── Performance
└── Security

Platform — use when compatibility scope materially changes effort:

Web              Mobile
├── Chrome        ├── Android
├── Firefox       └── iOS
├── Safari
└── Edge

Risk-based — use when time is constrained and effort must follow business risk:

Critical — Payment, Order creation
High     — Tax calculation, Shipping
Medium   — Email confirmation
Low      — Cosmetic layout
Caveat: these are not competing methodologies to pick one from. A real plan combines them — "Payment → API → external provider → critical risk → Chrome + mobile → regression + integration tests" describes one normal slice of scope, not an error.

Analogy and reference-class estimationCommon

Analogy estimates use similar past work as evidence. Reference-class thinking broadens that comparison to a class of comparable outcomes rather than relying on one memorable project.

Practical use: Choose comparable items by characteristics that drive testing effort: integration count, data complexity, platform count, regulatory constraints, novelty and dependency profile.

Caveat: Similarity by feature name is weak evidence. Record why the reference class is comparable and what is materially different.

Historical ratios and extrapolation from actualsCommon

Historical ratios can turn known quantities into a forecast, but only when the relationship is stable enough for the local context. Examples include test-design effort per comparable story class or review time per automation change.

Practical use: Use multiple completed samples, show the distribution or spread, and recalculate the ratio periodically. Prefer local actuals over generic industry percentages.

Caveat: A historical ratio is evidence from a past system. Architecture, team skill, tooling or definition-of-done changes can invalidate it.

One-time setup work vs recurring per-cycle workCommon

Initial setup and recurring execution have different cost structures. A framework change, environment bootstrap or data generator may be expensive once and cheap per cycle; manual regression may show the reverse pattern.

Practical use: Estimate one-time and recurring components separately, then state the number of cycles or releases assumed by any ROI comparison.

Caveat: Combining setup and recurring costs into one average hides the break-even point and makes automation ROI easy to misrepresent.

Summary

  • This chapter covers 10 required concepts while keeping tool/formula details tied to a practical decision.
  • Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
  • Claims that depend on a standard or product are grounded in the source registry below.

Source registry

Verified 16 Aug 2026
ISTQB Certified Tester Foundation Level Syllabus v4.0.1

Testing vocabulary, estimation techniques, coverage and test-management foundations

ISTQB · verified
Source ↗
ISTQB Glossary

Standard definitions for testing, defects, coverage and estimation terminology

ISTQB · verified
Source ↗
NASA Cost Estimating Handbook v4.0

WBS, analogy, uncertainty, risk, documenting assumptions and estimate calibration

NASA · verified
Source ↗
The Scrum Guide — November 2020

Scrum commitments, Product Backlog sizing and empiricism without prescribing story points

Ken Schwaber and Jeff Sutherland · verified
Source ↗