GimmeJob
Sign in
QA metrics & estimation · Chapter 08 / 08

QA Metrics & Estimation

Calibration & communication

A mature estimation practice learns. It keeps assumptions, retains prior forecasts, compares them with actual outcomes and changes the model when bias or structural change becomes visible.

Estimate/forecast + assumptions → decision → actual outcome → error/calibration review → model/process update → next estimate

Assumptions log, range and confidenceLess common

An assumptions log records the conditions that make an estimate or forecast valid. Pairing it with a range and confidence statement lets stakeholders see both the expected outcome and the reasons it may change.

Practical use: Keep assumptions testable where possible: environment ready by date, API contract stable, one reviewer available, migration dataset below a stated size.

Caveat: A long assumption list that nobody reviews is documentation debt. Link each important assumption to an owner or trigger.

Actual vs estimate and calibration over timeCommon

Calibration compares forecast claims with outcomes over repeated work. It can reveal optimism, pessimism, poor reference classes or models that are too narrow even when the average error looks acceptable.

Practical use: Track forecast error and whether outcomes fell inside stated ranges. Review by work class instead of blending unrelated tasks into one score.

Caveat: Do not turn estimate accuracy into an individual performance KPI; that encourages defensive padding and suppresses honest uncertainty.

Reforecasting as scope and evidence changeCommon

Reforecast when a material input changes or at a cadence appropriate to the decision horizon. The goal is to keep the decision model current, not to preserve the first number for appearance.

Practical use: Use explicit triggers and publish deltas: what changed, why the forecast moved, what new evidence is included and which decision is affected.

Caveat: A forecast that changes without explanation damages trust; a forecast that never changes despite new evidence is usually stale.

Communicating estimates and uncertainty to stakeholdersCommon

Good communication connects uncertainty to a decision. State the current range, major assumptions, confidence, what is driving the uncertainty, and what information would narrow it.

Practical use: Use scenario language: 'If environment X is ready Monday, 6–8 days; if migration defects continue at the current rate, 9–12.' Then state when the forecast will be refreshed.

Caveat: Do not bury uncertainty in a footnote while putting one precise date in the headline. The headline should represent the actual decision information.
Realistic QA planning example — Checkout / payment feature

One scenario, carried through every step of this course: decomposition, estimation, capacity, execution, KPI, quality gate, production, calibration. All numbers are a local scenario, not an industry standard.

Step 1 — Scope

Feature: New Checkout
Components: Web UI, Backend API, Payment provider,
Database, Confirmation email, Mobile web

Step 2 — Decomposition

Requirements review, Risk analysis, Test design,
Test data, Environment preparation, API testing,
UI testing, Payment integration, DB validation,
Mobile validation, Regression, Reporting

Step 3 — Risk (drives allocation, see risk-based decomposition above)

AreaRisk
PaymentCritical
Order creationCritical
Tax calculationHigh
EmailMedium
UI layoutMedium
Cosmetic stylingLow

Step 4 — Bottom-up estimate

Requirements review        4 h
Risk analysis               2 h
Test design                 8 h
Test data                   4 h
API testing                 8 h
UI testing                 10 h
Payment integration        10 h
DB validation                3 h
Mobile verification          4 h
Regression                   8 h
Reporting                    2 h
------------------------------
Base effort                 63 h

Step 5 — Three-point estimate

O = 48 h,  M = 64 h,  P = 96 h
PERT: E = (48 + 4×64 + 96) / 6 = 66.7 h

Range: 48–96 h
Weighted planning value: ~67 h

Keep the range — do not replace it with only "66.7 h".

Step 6 — Capacity

2 QA × 8 h × 5 days = 80 h theoretical

Meetings      8 h
Support       5 h
Other work    7 h
---------------
Real capacity 60 h

Expected effort ≈ 67 h vs real capacity 60 h

The five-day target is at risk. 80 theoretical hours exceeding 67 does not prove the work fits — only the 60 real hours count. Options: reduce scope, move low-risk coverage out, add capacity where parallelizable, or extend duration.

Step 7 — Execution evidence

Planned = 180, Executed = 170
Passed = 160, Failed = 8, Blocked = 2

Execution progress = 170/180 = 94.4%
Pass rate = 160/170 = 94.1%

Pass rate alone is not a release decision.

Step 8 — Quality gate

Critical defects = 0                  PASS
High defects = 2                      FAIL
Critical requirement coverage = 100%  PASS
Regression pass rate = 96%            FAIL

Quality gate = FAILED

Step 9 — Production metric

Pre-release confirmed defects = 23
Post-release confirmed defects = 2

Escape Rate = 2 / 25 × 100 = 8%

This denominator (23+2=25) is this scenario's explicit cohort, not a universal formula.

Step 10 — Calibration

Estimate = 67 h,  Actual = 74 h,  Variance = +7 h
Cause: payment sandbox instability

Learning: future payment-integration estimates should
explicitly model external sandbox availability/contingency.
Estimate → Actual → Difference → Cause → Model update

That closing loop — not the individual estimate accuracy of any one person — is the point of this whole chapter.

Summary

  • This chapter covers 4 required concepts while keeping tool/formula details tied to a practical decision.
  • Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
  • Claims that depend on a standard or product are grounded in the source registry below.

Source registry

Verified 16 Aug 2026
NASA Cost Estimating Handbook v4.0

WBS, analogy, uncertainty, risk, documenting assumptions and estimate calibration

NASA · verified
Source ↗
NIST/SEMATECH e-Handbook of Statistical Methods

Distributions, percentiles, variation and statistical interpretation

NIST · verified
Source ↗
ISTQB Certified Tester Foundation Level Syllabus v4.0.1

Testing vocabulary, estimation techniques, coverage and test-management foundations

ISTQB · verified
Source ↗