GimmeJob
Sign in
QA metrics & estimation · Chapter 02 / 08

QA Metrics & Estimation

QA & product quality metrics

This chapter is the metric catalog: every named defect and quality metric a QA Lead is actually asked to report, each with its formula, a worked scenario and the decision it should trigger. The framing is short on purpose — the metrics are the point.

customer / production evidence
          ↓
  classify + cohort + severity
          ↓
  defect / incident metrics
          ↓
  trend + root-cause grouping
          ↓
  prevention / release decision / gate

From raw measure to metric to KPI to target to actionCommon

Start from observable events, turn them into a defined metric, promote only decision-relevant metrics to KPI, and attach a target only when you have evidence for one.

Practical use: document the action path for each KPI — reviewer, trigger, next evidence checked, who approves corrective work.

Caveat: targets copied from another company are not baselines. Start from local history and business risk.

Define and track QA KPIs: escape rate, automation coverage, cycle time and MTTDCommon

Four different questions, not one mandatory scorecard: a customer-facing outcome (escape rate), a feedback-health signal (automation coverage), a delivery-flow signal (cycle time) and a detection signal (MTTD).

Practical use: define each separately, then read them together — rising escape rate with falling cycle time suggests faster delivery cut effective coverage.

Caveat: never average these into one quality number; they use different units and sit at different causal layers.
KPI exampleDefinition contractDecision it can support
Escape ratepost-gate defects / agreed defect cohort, per release windowstrengthen prevention or release controls
Automation coverageautomated eligible scope / eligible scopedecide where repetitive feedback is missing
Cycle timeelapsed time between defined workflow boundariesinvestigate queues and bottlenecks
MTTDincident start/first-observable time → detection timeimprove telemetry and alerting

Setting a target: baseline, scope and observation windowCommon

A baseline describes current behavior; a target describes the state worth moving toward, for a stated scope and window.

Practical use: derive initial targets from historical distribution, customer expectation and risk tolerance, and record when/why the target was set.

Caveat: a target with no stable definition invites denominator-shopping and cherry-picked windows.

Team and product metrics vs individual performance metricsCommon

Quality emerges from a socio-technical system — requirements, code, review, deployment, observability. Team/product metrics track that system better than individual counts do.

Practical use: use metrics to improve the system; keep any individual signal contextual, for coaching, never for ranking.

Caveat: bugs-found-per-tester, test-cases-written and tickets-closed reward volume and are trivially gamed — see the anti-pattern list this chapter builds toward.

Designing a balanced QA Lead scorecardCommon

A scorecard answers a small set of questions: are users harmed, are we detecting problems quickly, is feedback reliable, is delivery flowing, are known risks shrinking.

Practical use: one or two signals per question, trend not snapshot, each signal linked to an owner.

Caveat: a dashboard with twenty KPIs has no priorities. Keep diagnostics available but reserve the scorecard for decision-level signals.

Worked exercises: raw metric → KPI, missing denominator and vanity KPICommon

"12 production bugs" becomes useful only after you add a release cohort, severity policy, exposure window and denominator.

Practical use: take any raw number through three transformations — add a denominator, add a window, state the action a movement triggers. If you cannot state the action, it stays diagnostic telemetry.

Caveat: do not force every useful diagnostic signal into KPI status; a rich diagnostic layer and a small KPI set coexist fine.

Escaped defects, defect leakage and escape rateCommon

Escape rate is the percentage of defects first found after a chosen quality boundary (usually production or a release gate), against an agreed defect cohort.

Escape Rate = escaped defects / total defects in the agreed cohort × 100

Practical use: fix numerator, denominator, gate, severity handling, scope and window before comparing trend over time.

Caveat: there is no single industry denominator — some teams use all confirmed cohort defects, others compare pre- vs post-release only. Label the formula explicitly every time you report it.
scenario example — not a universal standard
numerator: defects first detected after the agreed gate
denominator: all confirmed defects in the agreed cohort
scope: release R42, Sev1–Sev3
window: 30 days after production release
target: none implied by this example; use a local baseline/risk decision
owner: QA Lead
decision: investigate prevention/release-control gaps when trend materially worsens

Defect Detection Percentage (DDP), the mirror of escape rateLess common

DDP is the percentage of all known cohort defects that were caught before release — the complement of escape rate over the same cohort.

DDP = pre-production defects / (pre-production + production defects) × 100

Practical use: report DDP alongside escape rate, not instead of it — DDP reads better to stakeholders ("92% caught" vs "8% escaped") but must share the exact same cohort definition or the two numbers silently diverge.

Caveat: DDP looks best when total defect volume is small. A DDP of 92% on 100 known defects is a very different signal from 92% on 8 known defects — check the sample size before trusting the percentage.

Defect removal efficiency (DRE)Less common

DRE is the share of a defect cohort removed before a downstream boundary, typically comparing pre-release removals against pre- plus post-release totals.

DRE = defects removed before release / (defects removed before + found after release) × 100

Practical use: use a closed or mature cohort so late-arriving production defects do not bias the denominator; keep severity and duplicate rules explicit.

Caveat: DRE is not product quality — a high value can coexist with a large total defect volume or poor user outcomes.
scenario definition — not a universal standard
numerator: confirmed defects removed before release
denominator: confirmed defects removed before + confirmed defects first found after release
scope: release cohort, agreed severities and duplicate policy
window: close after the agreed post-release observation period
target: none by default; establish from local baseline and risk appetite
owner: QA Lead with product/engineering
decision: investigate prevention effectiveness when the stable-definition trend worsens

Defect density and its limitationsLess common

Defect density divides confirmed defects by a size or exposure unit — KLOC, function points, component count or another locally meaningful denominator.

Defect Density = confirmed defects / size unit (e.g. KLOC, component count)

Practical use: use it only where the denominator is stable and comparable — within one product it can flag unusually defect-prone components worth investigating.

Caveat: density comparisons across languages, architectures or counting methods are usually invalid; more detected defects can mean better testing, not worse code.

Defect rejection rate: not-a-bug, duplicate and invalid reportsLess common

Defect rejection rate is the share of reported defects closed as not-a-bug, duplicate, expected behavior or invalid environment.

Rejection Rate = rejected reports / total reported defects × 100

Practical use: track it as a signal about report quality and requirement clarity, not tester performance — high rejection often points to ambiguous acceptance criteria or missing environment documentation, not bad testers.

Caveat: a target to minimize rejection rate can push testers to under-report edge cases they are unsure about. Pair it with a review of what got rejected, not just the count.

Defect severity distribution over raw defect countCommon

A distribution (critical/high/medium/low counts) is almost always more decision-relevant than a single total.

scenario example — not a universal standard
Critical: 2   High: 8   Medium: 24   Low: 50

Practical use: gate decisions on the severity distribution — specifically on open critical/high — not on the total. "23 defects, 0 critical, 2 high" and "23 defects, 6 critical" are different release conversations even though the totals match.

Caveat: severity labels drift between reporters and over time. Audit severity assignment periodically or the distribution itself becomes unreliable.

Defect aging and reopen rateCommon

Defect age is elapsed time in a defined state or from report to resolution. Reopen rate is how often a resolved defect returns to active status under agreed rules.

Reopen Rate = reopened defects / closed defects × 100

Practical use: segment age by severity and workflow state; for reopen rate use a closed-defect cohort, not all created defects.

Caveat: a low reopen rate can be artificial if teams clone defects instead of reopening them — audit workflow behavior before trusting the number.

Defect resolution time as a distribution, not one averageCommon

Resolution time is usually skewed — many defects close fast, a few sit for a long time — so a mean hides the tail that actually matters.

Practical use: show median plus a tail percentile or age bands (0–2 days, 3–7, 8–14, >14), segmented by severity.

Caveat: percentiles need adequate sample size; with a small cohort, show individual ages or bands instead of a false p95.

Customer-reported defects and production incident rateCommon

Direct signals of experienced harm, but raw counts are shaped by user volume, reporting channels and exposure.

Practical use: normalize by a meaningful exposure unit — active users, transactions, releases — when possible; keep raw support tickets separate from confirmed product defects until triage.

Caveat: a falling report count can mean better quality, fewer users, a broken support channel, or lower reporting propensity — cross-check with usage telemetry before celebrating.

Root-cause grouping and release-quality trend without a fake single scoreCommon

Release quality reads better as a profile — escaped defects, incidents, customer impact, rollback/hotfix behavior, recurring root-cause categories — than as one index.

Practical use: track recurring causes (requirement gaps, test-data gaps, environment mismatch, observability gaps) and use the trend to prioritize prevention work.

Caveat: do not collapse severity, volume and impact into an opaque composite score unless every component stays individually visible.

KPI vs KRI: how are we performing vs where is risk accumulatingCommon

A KPI answers how are we performing. A KRI — Key Risk Indicator — answers a different question: where is risk accumulating, before it has necessarily caused a failure yet.

KPI:
Defect Escape Rate
Current: 4.8%
Target: < local agreed threshold

KRI examples:
Untested critical requirements = 3
Open blocker defects = 2
Critical regression tests failing = 7
Untested external integrations = 1
Critical tests quarantined > 7 days = 4

Practical use: read them together, not as competing signals. "Production escape rate remained acceptable" (KPI) and "two critical payment paths are currently untested" (KRI) can both be true at the same time — that is exactly why the KRI is useful: it surfaces exposure the lagging KPI has not caught up to yet.

Caveat: a KRI without an owner or review cadence is just background noise. Give it the same contract a KPI gets: owner, review point, and the action a bad value triggers.

Quality gates: hard vs soft, and a worked release-gate exampleCommon

A quality gate is the bridge from metric to release decision:

Measure
   ↓
Metric
   ↓
Target / Threshold
   ↓
Quality Gate
   ↓
Decision
scenario example — not a universal standard
Measure: 2 open critical defects
Metric: Open critical defects = 2
Threshold: must be 0 for this release class
Quality gate: FAILED
Decision: release requires remediation or explicit risk acceptance

A hard gate blocks automatically (a critical security vulnerability blocks deployment, no exception). A soft gate raises a warning and routes to a human decision (automation pass rate just under target routes to QA Lead + Product Owner).

scenario example — not a universal standard
release gate:
  critical_defects: 0            (hard)
  high_defects: "<= 3"           (soft — review if exceeded)
  regression_pass_rate: ">= 98%" (hard)
  critical_requirement_coverage: "100%" (hard)
result if high_defects = 2, regression_pass_rate = 96%:
  high_defects            PASS
  regression_pass_rate    FAIL → gate blocked, escalate to release owner

Practical use: write the gate as executable-looking config, not prose, and log the pass/fail result of every dimension every release — not just the final verdict.

Not every metric needs to become a gate: metric ≠ target, and target ≠ gate. Pipeline runtime is useful but usually should not block a release by itself; a rising flaky-test trend can trigger improvement work without blocking every deployment; a critical security finding is exactly the kind of thing that should be a hard gate.

Caveat: a gate with too many hard conditions gets bypassed under deadline pressure until it is meaningless. Reserve hard gates for the few conditions the team will actually honor under pressure.
Published real-world example — Google App Engine moved HTTP status from logs to metrics

Google’s SRE Workbook gives a concrete production example. App Engine customers had HTTP status codes in logs, while the metrics view exposed only a global error rate. Diagnosing an error required finding the time on the graph, reading logs, and manually correlating the two.

The App Engine team exported HTTP status code as a bounded label on the request metric, illustrated in the SRE material as categories such as requests_total{status=404} and requests_total{status=500}. That let graphs separate error categories and allowed different alert behavior for client and server errors.

QA lesson: “error rate” can be too aggregated to support a decision. A small, controlled dimension such as HTTP status class/code can turn a vanity-looking total into diagnostic evidence. The example also shows why dimensions need bounded cardinality rather than arbitrary user IDs or request values.

Google SRE Workbook — Monitoring, real-world examples

Summary

  • This chapter covers 18 required concepts while keeping tool/formula details tied to a practical decision.
  • Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
  • Claims that depend on a standard or product are grounded in the source registry below.

Source registry

Verified 16 Aug 2026
ISTQB Certified Tester Foundation Level Syllabus v4.0.1

Testing vocabulary, estimation techniques, coverage and test-management foundations

ISTQB · verified
Source ↗
ISTQB Glossary

Standard definitions for testing, defects, coverage and estimation terminology

ISTQB · verified
Source ↗
ISO/IEC 25010:2023 — Product quality model

Product-quality characteristics used to define what quality signals should measure

ISO/IEC · verified
Source ↗
NIST/SEMATECH e-Handbook of Statistical Methods

Distributions, percentiles, variation and statistical interpretation

NIST · verified
Source ↗