GimmeJob
Sign in
QA metrics & estimation · Chapter 03 / 08

QA Metrics & Estimation

Test execution & automation metrics

Execution and automation numbers measure the feedback system, not the product. This chapter walks through each one — execution status, coverage, and automation-health signals — with a formula first and the decision it supports second.

planned risk / test basis
        ↓
 selected execution scope
        ↓
 run results ──→ pass / fail / blocked
        ↓
 feedback-system health
        ├── stability / flakiness / false failures
        ├── coverage denominator
        ├── runtime / maintenance cost
        └── diagnosis time

Planned, executed, blocked and pass/fail statusCommon

Four counts, four different questions, all needing the same stable denominator.

Execution Progress = executed tests / planned tests × 100
Pass Rate = passed tests / executed tests × 100
Fail Rate = failed tests / executed tests × 100
Blocked Rate = blocked tests / executed tests × 100

Practical use: report completion and outcome together — executed/planned plus passed/failed/blocked among executed. Keep blocked separate; it usually signals environment, data or dependency risk, not test failure.

Caveat: never treat unexecuted tests as passes. A 100% pass rate on 20% of planned scope is not release evidence for the missing 80%.

Why pass rate is not a standalone quality KPICommon

Pass rate is the share of executed checks that passed under one run. It says nothing about whether the suite covers the risks that matter or whether users are harmed elsewhere.

Practical use: read pass rate together with execution coverage, test relevance and production evidence — never alone.

Caveat: a weak suite can post 100% pass rate every day; a strong suite may show a lower rate precisely because it caught a real regression.

Requirement coverage and risk-weighted coverageCommon

Requirement Coverage = requirements with mapped tests / total requirements × 100
Risk Coverage = high-risk items with test coverage / total high-risk items × 100

Practical use: state the coverage criterion — mapped, designed, automated, executed or passed are different claims. For risk-weighted views keep the underlying risk items visible, not just the ratio.

Caveat: a high percentage can hide an incomplete test basis. Coverage never proves the absence of defects.

Code coverage and mutation score as specialist signals, not proofCommon

Code coverage (statement, branch, function, line) measures which structural elements executed; mutation score measures whether tests would catch a deliberately introduced code change.

Practical use: use coverage to find unexercised code and mutation score to challenge assertion strength where the cost is justified — inspect surviving mutants rather than chasing a target number.

Caveat: high coverage coexists easily with weak assertions or missing scenarios. Evidence about the test suite is not proof of product quality.

Regression completion without a universal percentage targetCommon

Regression completion means the agreed scope reached a defined execution state — a scope sized to change impact and risk, not a "run 100% every time" rule.

Practical use: report what was selected and why, what remains untested, and whether blocked items touch critical risk. A smaller risk-based suite often gives better evidence than a large low-value one.

Caveat: completion is not readiness. Release decisions also need defect status, residual risk and relevant non-functional evidence.

Automation health: the governance layer above framework metricsCommon

Automation health asks whether automated feedback is trustworthy, timely and maintainable — the layer above framework-specific implementation metrics.

Practical use: review flakiness, false-failure rate, coverage denominators, runtime and diagnosis time together as one health view, not five disconnected charts.

Caveat: a suite with thousands of green tests is not healthy if teams ignore failures or need repeated reruns to get green.

Flaky-test rate, retry rate and quarantine backlog/ageCommon

Flaky Rate = tests classified flaky / tests observed under the classification rule × 100
Retry Rate = executions requiring a retry / total executions in scope × 100

Practical use: define whether flakiness is measured per test, per execution or per incident. Track quarantined tests with owner, reason and age so quarantine stays temporary containment, not a graveyard.

Caveat: retries can mask real product defects as well as test defects — preserve first-attempt outcomes and investigate before classifying.
scenario definitions — not universal standards
flaky-rate numerator: tests classified flaky
flaky-rate denominator: tests observed under the agreed classification rule
retry-rate numerator: executions requiring retry
retry-rate denominator: total executions in scope
scope: merge-gate suite
window: rolling 14 days
target: none implied; baseline first, then set a local guardrail if useful
owner: automation maintainers / QA lead
decision: prioritize reliability work or quarantine removal when feedback trust degrades

False failure rate: infrastructure noise vs real regressionsLess common

A false failure fails for a reason unrelated to the product under test — infra outage, network blip, test-data collision, environment misconfiguration. It is a distinct category from a flaky test (which fails inconsistently for no clear reason) and from a real regression.

False Failure Rate = failures classified as infra/test-code cause / total failures × 100

Practical use: triage every failure into one of three buckets — real regression, flaky, false failure — before reporting a pass rate. A rising false-failure rate points at CI/environment investment, not at the product or the test design.

Caveat: classifying a failure as "false" without investigation is how real regressions get waved through. Require a documented reason, not a gut call, before excluding a failure from the product-quality signal.

Automation coverage denominators, in depthCommon

"Automation coverage" has no meaning without a stated denominator — eligible regression scenarios, API endpoints, critical journeys, requirements, risk items or platform combinations are all valid but different choices.

Automation Coverage = automated eligible scope / eligible scope × 100

Practical use: pick the denominator tied to the actual decision — merge-gate eligible scope for feedback speed, critical-journey scope for business protection.

Caveat: automating low-value cases inflates the percentage while important risk stays manual or uncovered. Show excluded scope and the reason.

Merge-gate runtime and time to diagnoseCommon

Pipeline runtime measures feedback latency; time to diagnose measures how long an engineer needs to understand a failure well enough to act. Both affect developer flow, not just QA.

Practical use: measure critical-path runtime separately from total parallel compute time; for diagnosis time, define start (first failure notification) and end (cause category identified) events precisely.

Caveat: a faster pipeline that drops useful checks can be worse than a slower, more complete one. Read runtime together with coverage and false-failure rate.

Automation execution-time trend and maintenance costLess common

Two secondary but real signals: how the suite's wall-clock runtime moves over time, and how much human effort it costs to keep the suite green.

scenario example — not a universal standard
January regression runtime: 120 min
February: 95 min
March: 70 min
maintenance cost (same window): engineer-hours/sprint spent on repair, or average repair time per failed test

Practical use: track runtime as a trend, not a single value, and pair it with maintenance cost — a suite that got faster by deleting flaky coverage is not actually healthier.

Caveat: falling runtime alone is not an achievement. Confirm coverage and false-failure rate held steady before crediting the pipeline change.

Automation stability vs pass rateCommon

Stability is about the repeatability of the signal; pass rate is about whether expectations were met. A deterministic failing test is stable even though it is red.

Practical use: separate product failures, deterministic test defects, infrastructure failures and flaky outcomes in reporting — this is what makes remediation ownership clear.

Caveat: labeling every intermittent failure "flaky" can hide a real environment incident or race condition in the product itself.
Real GimmeJob observation — one CI + deploy run on 16 Aug 2026

A real GimmeJob GitHub Actions run provides a useful example of the difference between an observation and a metric baseline.

Observed run:

  • Workflow: CI and Cloudflare deploy
  • Branch: main
  • Commit: 0c71ba630fe5dc2e08e3194cdc79067b6548c894
  • Trigger: push
  • Conclusion: success
  • Started/created: 2026-08-16T14:08:49Z
  • Updated/completed: 2026-08-16T14:12:21Z
  • Observed wall-clock interval: 3 min 32 s

Open the real GitHub Actions run

What this proves: that specific run completed successfully and the recorded run interval was 3 min 32 s.

What this does not prove: that 3 min 32 s is the normal pipeline duration, a p95, a target, an SLO, or evidence that product quality is good. One sample cannot establish a distribution.

Turn it into a useful metric: collect comparable main runs using the same event boundary, report median plus a tail percentile when sample size is sufficient, and segment failures/retries or major workflow changes instead of hiding them in one average.

Summary

  • This chapter covers 12 required concepts while keeping tool/formula details tied to a practical decision.
  • Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
  • Claims that depend on a standard or product are grounded in the source registry below.

Source registry

Verified 16 Aug 2026
ISTQB Certified Tester Foundation Level Syllabus v4.0.1

Testing vocabulary, estimation techniques, coverage and test-management foundations

ISTQB · verified
Source ↗
ISTQB Glossary

Standard definitions for testing, defects, coverage and estimation terminology

ISTQB · verified
Source ↗
DORA software delivery performance metrics

Current five software-delivery performance metrics and context-sensitive interpretation

Google Cloud / DORA · verified
Source ↗
NIST/SEMATECH e-Handbook of Statistical Methods

Distributions, percentiles, variation and statistical interpretation

NIST · verified
Source ↗