GimmeJob
Sign in
QA metrics & estimation · Chapter 04 / 08

QA Metrics & Estimation

Delivery & production metrics

Delivery and production metrics answer adjacent but different questions: how work flows, how deployments perform, and how users experience the running service. Keeping those boundaries clear prevents dashboard vocabulary from becoming meaningless.

work flow ──→ deployment ──→ running service ──→ user impact
   │              │                │                  │
cycle/WIP      DORA metrics      SLI/SLO          incidents
throughput     change failure    latency          customer harm

Lead time, cycle time, throughput and WIPCommon

Flow metrics describe how work moves through a defined system. Cycle time measures elapsed time from a chosen start to finish, throughput counts finished items per unit time, and WIP counts started but unfinished items. Lead time may use an earlier customer/request boundary, so teams must define it locally.

Practical use: Use the Kanban flow definitions consistently and visualize distributions rather than only monthly averages. Little’s Law can be a useful consistency relationship when the system is sufficiently stable.

Caveat: Changing workflow boundaries changes the metric. Do not compare teams unless their definitions and work-item granularity are comparable.
Little's Law (steady-state relationship)
average WIP ≈ average throughput × average cycle time

Use it as a system check, not as a promise for one individual item.

Work-item age, queue time and blocked timeLess common

Work-item age is elapsed time since an unfinished item started. Queue time isolates waiting; blocked time records periods when progress cannot continue because a dependency or condition is unresolved.

Practical use: Use age to surface currently stuck work before it becomes long cycle time. Track reasons for blocking to target systemic constraints rather than individual pressure.

Caveat: Teams often record blocked state inconsistently. If the workflow does not capture it reliably, treat the metric as incomplete evidence.

Batch size and release frequencyCommon

Batch size describes how much change moves together through a boundary; release or deployment frequency describes how often changes cross that boundary. Smaller batches can reduce risk and shorten feedback loops when the delivery system supports them.

Practical use: Measure the boundary that matters: merge, deployment or end-user release are not the same. Pair frequency with change-failure and customer-impact signals.

Caveat: Higher frequency is not automatically better if it is achieved with unstable deployments, hidden rework or ineffective controls.

Why velocity and story points should not compare teamsCommon

Story points are a local relative sizing convention, not a standardized unit. The Scrum Guide requires transparency and a Product Backlog with enough information to plan work, but it does not prescribe story points as a universal unit.

Practical use: Use a team’s historical sizing only inside the context that created it. For forecasting across teams, prefer observable flow data or explicitly calibrated common units.

Caveat: Comparing velocity creates pressure to inflate estimates and destroys the meaning of the local scale.

The current five DORA metrics, in depthCommon

DORA currently defines five software-delivery performance metrics. Throughput is represented by change lead time, deployment frequency and failed deployment recovery time; instability is represented by change fail rate and deployment rework rate.

Practical use: Measure them for one application or service in a stable context and trend them over time. Keep DORA’s deployment-oriented definitions separate from generic incident MTTR or product-defect metrics.

Caveat: Old material may still describe the “Four Keys” or use MTTR. Treat the current DORA documentation as authoritative for the current model.
DORA factorCurrent metric
ThroughputChange lead time
ThroughputDeployment frequency
ThroughputFailed deployment recovery time
InstabilityChange fail rate
InstabilityDeployment rework rate

Application context, trend over ranking, metrics as signals not goalsCommon

DORA explicitly emphasizes context and measuring an application or service over time. The metrics are most useful for identifying improvement opportunities and validating change.

Practical use: Use a baseline, observe changes after an intervention and inspect all five metrics together. Explain architecture, regulatory and release context before external comparisons.

Caveat: Turning DORA into individual targets or a league table invites gaming and ignores the research model’s context.

Where DORA connects to deeper reliability materialLess common

DORA measures delivery outcomes. Observability and SRE explain how services expose behavior, detect user-impacting failures and manage reliability objectives.

Practical use: Use DORA to ask whether delivery is fast and stable; use SLI/SLO, telemetry and incident analysis to understand runtime reliability in more depth.

Caveat: Do not label every operational metric “DORA.” Availability, latency and MTTD are useful but belong to other measurement models.

Availability, error/success rate, latency percentiles and crash-free sessionsCommon

Production-quality signals should reflect user-visible service behavior. Availability and success rate describe whether operations succeed; latency percentiles show response-time distribution; crash-free sessions/users are common mobile stability views.

Practical use: Define the SLI at the point closest to the user experience and state exclusions explicitly. Use percentiles for latency and segment critical journeys where aggregate availability hides important failures.

Caveat: A health endpoint can be green while user requests fail. Infrastructure health is a diagnostic input, not a substitute for user-facing SLIs.

MTTD and generic MTTR terminologyCommon

MTTD usually measures time from incident start or first observable harmful condition to detection. MTTR is overloaded: organizations use it for mean time to repair, recover, restore or resolve, which are not identical.

Practical use: Name the exact event boundaries instead of relying on the acronym. Record whether timestamps are automatic or manually inferred.

Caveat: DORA’s failed deployment recovery time is a specific deployment metric and should not be silently renamed generic MTTR.

Incident and customer-impact signals for QA decisionsCommon

QA can use production incidents as feedback about missed risks, detection gaps and release controls. Useful fields include affected users/transactions, duration, severity, change linkage and detection path.

Practical use: Connect incidents back to test/risk models and prevention actions. Trend cause categories and recurrence, not only incident counts.

Caveat: Incident volume is exposure-sensitive and severity-sensitive. One severe incident can matter more than many minor ones.

SLI, SLO and error-budget vocabularyLess common

An SLI is a quantitative measure of service behavior; an SLO is a target range or objective for that indicator over a window; an error budget is the allowed unreliability implied by the objective.

Practical use: Use the vocabulary to connect quality decisions with operational reliability. Detailed SLO design, burn rates and telemetry implementation belong in the Observability & SRE path.

Caveat: An SLO is not a universal target. It must reflect user expectations, business consequences and system capability.
Real boundary check — why a successful CI run is not automatically a DORA metric

DORA’s current model defines change lead time from commit to successfully running in production, deployment frequency around production deployments/releases, change fail rate around deployments/releases that require remediation, failed deployment recovery time after failed deployments, and deployment rework rate for unplanned deployments caused by production incidents/user-facing bugs.

Now compare that contract with the real GimmeJob Actions example in the previous chapter. A GitHub Actions run duration can be a valid CI feedback-time observation. It only contributes to a DORA production metric when the events you capture actually satisfy DORA’s production/release boundary. Calling every pipeline duration “DORA lead time” would silently change the definition.

Practical implementation: store separate timestamps for commit, deployment start, successful production deployment/release, failure detection, remediation deployment, and recovery. Derive each metric from those events instead of renaming a convenient CI timer.

DORA software delivery performance metrics · DORA Quick Check

Summary

  • This chapter covers 11 required concepts while keeping tool/formula details tied to a practical decision.
  • Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
  • Claims that depend on a standard or product are grounded in the source registry below.

Source registry

Verified 16 Aug 2026
The Kanban Guide

WIP, throughput, work item age and cycle time definitions for flow measurement

Kanban Guides / ProKanban.org · verified
Source ↗
DORA software delivery performance metrics

Current five software-delivery performance metrics and context-sensitive interpretation

Google Cloud / DORA · verified
Source ↗
A history of DORA’s software delivery metrics

Evolution from the original four keys to the current five-metric model

Google Cloud / DORA · verified
Source ↗
Google Site Reliability Engineering books

SLIs, SLOs, error budgets, incident response and reliability measurement

Google · verified
Source ↗
OpenTelemetry concepts

Telemetry signals that support production-quality and detection metrics

OpenTelemetry · verified
Source ↗
Prometheus alerting practices

Actionable symptom-based alerting and operational detection signals

Prometheus · verified
Source ↗
Configure liveness, readiness and startup probes

Concrete production-detection signals and the difference between health checks

Kubernetes · verified
Source ↗