Test execution & automation metrics
Execution and automation numbers measure the feedback system, not the product. This chapter walks through each one — execution status, coverage, and automation-health signals — with a formula first and the decision it supports second.
planned risk / test basis
↓
selected execution scope
↓
run results ──→ pass / fail / blocked
↓
feedback-system health
├── stability / flakiness / false failures
├── coverage denominator
├── runtime / maintenance cost
└── diagnosis timePlanned, executed, blocked and pass/fail statusCommon
Four counts, four different questions, all needing the same stable denominator.
Execution Progress = executed tests / planned tests × 100
Pass Rate = passed tests / executed tests × 100
Fail Rate = failed tests / executed tests × 100
Blocked Rate = blocked tests / executed tests × 100Practical use: report completion and outcome together — executed/planned plus passed/failed/blocked among executed. Keep blocked separate; it usually signals environment, data or dependency risk, not test failure.
Caveat: never treat unexecuted tests as passes. A 100% pass rate on 20% of planned scope is not release evidence for the missing 80%.
Why pass rate is not a standalone quality KPICommon
Pass rate is the share of executed checks that passed under one run. It says nothing about whether the suite covers the risks that matter or whether users are harmed elsewhere.
Practical use: read pass rate together with execution coverage, test relevance and production evidence — never alone.
Caveat: a weak suite can post 100% pass rate every day; a strong suite may show a lower rate precisely because it caught a real regression.
Requirement coverage and risk-weighted coverageCommon
Requirement Coverage = requirements with mapped tests / total requirements × 100
Risk Coverage = high-risk items with test coverage / total high-risk items × 100Practical use: state the coverage criterion — mapped, designed, automated, executed or passed are different claims. For risk-weighted views keep the underlying risk items visible, not just the ratio.
Caveat: a high percentage can hide an incomplete test basis. Coverage never proves the absence of defects.
Code coverage and mutation score as specialist signals, not proofCommon
Code coverage (statement, branch, function, line) measures which structural elements executed; mutation score measures whether tests would catch a deliberately introduced code change.
Practical use: use coverage to find unexercised code and mutation score to challenge assertion strength where the cost is justified — inspect surviving mutants rather than chasing a target number.
Caveat: high coverage coexists easily with weak assertions or missing scenarios. Evidence about the test suite is not proof of product quality.
Regression completion without a universal percentage targetCommon
Regression completion means the agreed scope reached a defined execution state — a scope sized to change impact and risk, not a "run 100% every time" rule.
Practical use: report what was selected and why, what remains untested, and whether blocked items touch critical risk. A smaller risk-based suite often gives better evidence than a large low-value one.
Caveat: completion is not readiness. Release decisions also need defect status, residual risk and relevant non-functional evidence.
Automation health: the governance layer above framework metricsCommon
Automation health asks whether automated feedback is trustworthy, timely and maintainable — the layer above framework-specific implementation metrics.
Practical use: review flakiness, false-failure rate, coverage denominators, runtime and diagnosis time together as one health view, not five disconnected charts.
Caveat: a suite with thousands of green tests is not healthy if teams ignore failures or need repeated reruns to get green.
Flaky-test rate, retry rate and quarantine backlog/ageCommon
Flaky Rate = tests classified flaky / tests observed under the classification rule × 100
Retry Rate = executions requiring a retry / total executions in scope × 100Practical use: define whether flakiness is measured per test, per execution or per incident. Track quarantined tests with owner, reason and age so quarantine stays temporary containment, not a graveyard.
Caveat: retries can mask real product defects as well as test defects — preserve first-attempt outcomes and investigate before classifying.
scenario definitions — not universal standards
flaky-rate numerator: tests classified flaky
flaky-rate denominator: tests observed under the agreed classification rule
retry-rate numerator: executions requiring retry
retry-rate denominator: total executions in scope
scope: merge-gate suite
window: rolling 14 days
target: none implied; baseline first, then set a local guardrail if useful
owner: automation maintainers / QA lead
decision: prioritize reliability work or quarantine removal when feedback trust degradesFalse failure rate: infrastructure noise vs real regressionsLess common
A false failure fails for a reason unrelated to the product under test — infra outage, network blip, test-data collision, environment misconfiguration. It is a distinct category from a flaky test (which fails inconsistently for no clear reason) and from a real regression.
False Failure Rate = failures classified as infra/test-code cause / total failures × 100Practical use: triage every failure into one of three buckets — real regression, flaky, false failure — before reporting a pass rate. A rising false-failure rate points at CI/environment investment, not at the product or the test design.
Caveat: classifying a failure as "false" without investigation is how real regressions get waved through. Require a documented reason, not a gut call, before excluding a failure from the product-quality signal.
Automation coverage denominators, in depthCommon
"Automation coverage" has no meaning without a stated denominator — eligible regression scenarios, API endpoints, critical journeys, requirements, risk items or platform combinations are all valid but different choices.
Automation Coverage = automated eligible scope / eligible scope × 100Practical use: pick the denominator tied to the actual decision — merge-gate eligible scope for feedback speed, critical-journey scope for business protection.
Caveat: automating low-value cases inflates the percentage while important risk stays manual or uncovered. Show excluded scope and the reason.
Merge-gate runtime and time to diagnoseCommon
Pipeline runtime measures feedback latency; time to diagnose measures how long an engineer needs to understand a failure well enough to act. Both affect developer flow, not just QA.
Practical use: measure critical-path runtime separately from total parallel compute time; for diagnosis time, define start (first failure notification) and end (cause category identified) events precisely.
Caveat: a faster pipeline that drops useful checks can be worse than a slower, more complete one. Read runtime together with coverage and false-failure rate.
Automation execution-time trend and maintenance costLess common
Two secondary but real signals: how the suite's wall-clock runtime moves over time, and how much human effort it costs to keep the suite green.
scenario example — not a universal standard
January regression runtime: 120 min
February: 95 min
March: 70 min
maintenance cost (same window): engineer-hours/sprint spent on repair, or average repair time per failed testPractical use: track runtime as a trend, not a single value, and pair it with maintenance cost — a suite that got faster by deleting flaky coverage is not actually healthier.
Caveat: falling runtime alone is not an achievement. Confirm coverage and false-failure rate held steady before crediting the pipeline change.
Automation stability vs pass rateCommon
Stability is about the repeatability of the signal; pass rate is about whether expectations were met. A deterministic failing test is stable even though it is red.
Practical use: separate product failures, deterministic test defects, infrastructure failures and flaky outcomes in reporting — this is what makes remediation ownership clear.
Caveat: labeling every intermittent failure "flaky" can hide a real environment incident or race condition in the product itself.
Summary
- This chapter covers 12 required concepts while keeping tool/formula details tied to a practical decision.
- Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
- Claims that depend on a standard or product are grounded in the source registry below.