Delivery & production metrics
Delivery and production metrics answer adjacent but different questions: how work flows, how deployments perform, and how users experience the running service. Keeping those boundaries clear prevents dashboard vocabulary from becoming meaningless.
work flow ──→ deployment ──→ running service ──→ user impact
│ │ │ │
cycle/WIP DORA metrics SLI/SLO incidents
throughput change failure latency customer harmLead time, cycle time, throughput and WIPCommon
Flow metrics describe how work moves through a defined system. Cycle time measures elapsed time from a chosen start to finish, throughput counts finished items per unit time, and WIP counts started but unfinished items. Lead time may use an earlier customer/request boundary, so teams must define it locally.
Practical use: Use the Kanban flow definitions consistently and visualize distributions rather than only monthly averages. Little’s Law can be a useful consistency relationship when the system is sufficiently stable.
Caveat: Changing workflow boundaries changes the metric. Do not compare teams unless their definitions and work-item granularity are comparable.
Little's Law (steady-state relationship)
average WIP ≈ average throughput × average cycle timeUse it as a system check, not as a promise for one individual item.
Work-item age, queue time and blocked timeLess common
Work-item age is elapsed time since an unfinished item started. Queue time isolates waiting; blocked time records periods when progress cannot continue because a dependency or condition is unresolved.
Practical use: Use age to surface currently stuck work before it becomes long cycle time. Track reasons for blocking to target systemic constraints rather than individual pressure.
Caveat: Teams often record blocked state inconsistently. If the workflow does not capture it reliably, treat the metric as incomplete evidence.
Batch size and release frequencyCommon
Batch size describes how much change moves together through a boundary; release or deployment frequency describes how often changes cross that boundary. Smaller batches can reduce risk and shorten feedback loops when the delivery system supports them.
Practical use: Measure the boundary that matters: merge, deployment or end-user release are not the same. Pair frequency with change-failure and customer-impact signals.
Caveat: Higher frequency is not automatically better if it is achieved with unstable deployments, hidden rework or ineffective controls.
Why velocity and story points should not compare teamsCommon
Story points are a local relative sizing convention, not a standardized unit. The Scrum Guide requires transparency and a Product Backlog with enough information to plan work, but it does not prescribe story points as a universal unit.
Practical use: Use a team’s historical sizing only inside the context that created it. For forecasting across teams, prefer observable flow data or explicitly calibrated common units.
Caveat: Comparing velocity creates pressure to inflate estimates and destroys the meaning of the local scale.
The current five DORA metrics, in depthCommon
DORA currently defines five software-delivery performance metrics. Throughput is represented by change lead time, deployment frequency and failed deployment recovery time; instability is represented by change fail rate and deployment rework rate.
Practical use: Measure them for one application or service in a stable context and trend them over time. Keep DORA’s deployment-oriented definitions separate from generic incident MTTR or product-defect metrics.
Caveat: Old material may still describe the “Four Keys” or use MTTR. Treat the current DORA documentation as authoritative for the current model.
| DORA factor | Current metric |
|---|---|
| Throughput | Change lead time |
| Throughput | Deployment frequency |
| Throughput | Failed deployment recovery time |
| Instability | Change fail rate |
| Instability | Deployment rework rate |
Application context, trend over ranking, metrics as signals not goalsCommon
DORA explicitly emphasizes context and measuring an application or service over time. The metrics are most useful for identifying improvement opportunities and validating change.
Practical use: Use a baseline, observe changes after an intervention and inspect all five metrics together. Explain architecture, regulatory and release context before external comparisons.
Caveat: Turning DORA into individual targets or a league table invites gaming and ignores the research model’s context.
Where DORA connects to deeper reliability materialLess common
DORA measures delivery outcomes. Observability and SRE explain how services expose behavior, detect user-impacting failures and manage reliability objectives.
Practical use: Use DORA to ask whether delivery is fast and stable; use SLI/SLO, telemetry and incident analysis to understand runtime reliability in more depth.
Caveat: Do not label every operational metric “DORA.” Availability, latency and MTTD are useful but belong to other measurement models.
Availability, error/success rate, latency percentiles and crash-free sessionsCommon
Production-quality signals should reflect user-visible service behavior. Availability and success rate describe whether operations succeed; latency percentiles show response-time distribution; crash-free sessions/users are common mobile stability views.
Practical use: Define the SLI at the point closest to the user experience and state exclusions explicitly. Use percentiles for latency and segment critical journeys where aggregate availability hides important failures.
Caveat: A health endpoint can be green while user requests fail. Infrastructure health is a diagnostic input, not a substitute for user-facing SLIs.
MTTD and generic MTTR terminologyCommon
MTTD usually measures time from incident start or first observable harmful condition to detection. MTTR is overloaded: organizations use it for mean time to repair, recover, restore or resolve, which are not identical.
Practical use: Name the exact event boundaries instead of relying on the acronym. Record whether timestamps are automatic or manually inferred.
Caveat: DORA’s failed deployment recovery time is a specific deployment metric and should not be silently renamed generic MTTR.
Incident and customer-impact signals for QA decisionsCommon
QA can use production incidents as feedback about missed risks, detection gaps and release controls. Useful fields include affected users/transactions, duration, severity, change linkage and detection path.
Practical use: Connect incidents back to test/risk models and prevention actions. Trend cause categories and recurrence, not only incident counts.
Caveat: Incident volume is exposure-sensitive and severity-sensitive. One severe incident can matter more than many minor ones.
SLI, SLO and error-budget vocabularyLess common
An SLI is a quantitative measure of service behavior; an SLO is a target range or objective for that indicator over a window; an error budget is the allowed unreliability implied by the objective.
Practical use: Use the vocabulary to connect quality decisions with operational reliability. Detailed SLO design, burn rates and telemetry implementation belong in the Observability & SRE path.
Caveat: An SLO is not a universal target. It must reflect user expectations, business consequences and system capability.
Summary
- This chapter covers 11 required concepts while keeping tool/formula details tied to a practical decision.
- Definitions, scope, assumptions and caveats matter more than a number or a tool name by itself.
- Claims that depend on a standard or product are grounded in the source registry below.