Metrics, percentiles and objectives
A performance test produces several kinds of evidence. Read them together: user-visible outcomes tell you that behavior changed; resource and saturation metrics help explain the change.
Response time and latency
Terminology varies between tools, so always state the measurement boundary. For an HTTP test, the most useful number is often the full request duration as seen by the client. Other tools may separately expose connect time, TLS time, waiting/time-to-first-byte and receive time.
Do not compare two metrics merely because both are labelled “latency.” Compare what they actually measure.
Why average is insufficient
An average collapses the distribution into one number. It can remain healthy even when a small but important portion of requests becomes very slow.
A percentile describes a boundary in the observed distribution:
- p50: half the observations completed at or below this value;
- p95: 95% completed at or below it;
- p99: 99% completed at or below it.
If p95 is 300 ms and p99 is 2 s, most traffic is fast but the slow tail is materially worse. That information is lost in a single mean.
Percentiles still need context: workload, endpoint grouping, errors, duration and sample count. A p99 from a tiny sample is not a stable tail estimate.
Throughput
Throughput measures completed work per unit time. Depending on the goal, report:
- HTTP requests per second;
- business transactions per second/minute;
- messages processed per second;
- bytes/records/jobs per unit time.
A plateau is important. If offered load or concurrency rises while useful throughput stops rising and latency increases, the system is likely approaching a capacity constraint.
Errors and correctness
A fast 500 response is not a successful performance result. Performance scripts should validate enough functional behavior to prove that the measured transaction is actually correct.
Separate:
- protocol failures/timeouts;
- explicit non-success responses;
- assertion/check failures where the protocol response is technically successful but business output is wrong.
Under overload, error shape also matters: controlled 429 or queue rejection may be preferable to timeouts, corrupted state or cascading failures.
Resource and saturation metrics
Common server-side signals include:
- CPU utilization and throttling;
- memory, heap, garbage collection and allocation rate;
- thread/worker/connection pools and wait time;
- queue depth and age;
- database connections, locks, waits, query latency and I/O;
- disk latency/IOPS and network saturation;
- cache hit ratio and eviction;
- downstream dependency latency/errors;
- autoscaling and container/VM limits.
Choose from the architecture. A checklist that ignores the actual system is not observability.
Baseline
A baseline records a known system/version under controlled conditions so future runs can be compared. Keep the variables stable enough that a difference is interpretable:
- same workload model;
- comparable environment and data;
- same measurement definitions;
- known build/configuration;
- enough repetition to understand normal run-to-run variance.
Thresholds and SLOs
Grafana k6 thresholds turn metrics into pass/fail criteria. A representative example is:
export const options = {
thresholds: {
http_req_failed: ['rate<0.01'],
http_req_duration: ['p(95)<500', 'p(99)<1000'],
},
};These numbers are examples of syntax, not universal performance requirements. Replace them with real product objectives.
The value of a threshold is automation: a CI job can fail without a person reading every chart. The limitation is also important: a pass/fail result cannot diagnose the reason for a regression, and a noisy shared environment can create misleading gates. Preserve raw/trended evidence and telemetry.
Correlation is the analysis technique
Suppose p95 rises at minute 8 while throughput plateaus. At the same time:
- DB connection pool reaches its limit;
- pool wait time increases;
- CPU remains moderate;
- downstream latency is unchanged.
That time-aligned evidence supports a connection-pool/database bottleneck hypothesis. It is far stronger than saying “CPU was high somewhere during the test.”
Reporting rule
Every performance conclusion should name the workload and the relevant metric. “p95 = 420 ms” is incomplete. “At the defined 700 RPS checkout workload, p95 was 420 ms, errors 0.2%, throughput remained stable, and the database pool stayed below saturation” is actionable evidence.