GimmeJob
Sign in
Performance testing · Chapter 03 / 09

Performance Testing

Metrics, percentiles and objectives

A performance test produces several kinds of evidence. Read them together: user-visible outcomes tell you that behavior changed; resource and saturation metrics help explain the change.

Response time and latency

Terminology varies between tools, so always state the measurement boundary. For an HTTP test, the most useful number is often the full request duration as seen by the client. Other tools may separately expose connect time, TLS time, waiting/time-to-first-byte and receive time.

Do not compare two metrics merely because both are labelled “latency.” Compare what they actually measure.

Why average is insufficient

An average collapses the distribution into one number. It can remain healthy even when a small but important portion of requests becomes very slow.

A percentile describes a boundary in the observed distribution:

  • p50: half the observations completed at or below this value;
  • p95: 95% completed at or below it;
  • p99: 99% completed at or below it.

If p95 is 300 ms and p99 is 2 s, most traffic is fast but the slow tail is materially worse. That information is lost in a single mean.

Percentiles still need context: workload, endpoint grouping, errors, duration and sample count. A p99 from a tiny sample is not a stable tail estimate.

Throughput

Throughput measures completed work per unit time. Depending on the goal, report:

  • HTTP requests per second;
  • business transactions per second/minute;
  • messages processed per second;
  • bytes/records/jobs per unit time.

A plateau is important. If offered load or concurrency rises while useful throughput stops rising and latency increases, the system is likely approaching a capacity constraint.

Errors and correctness

A fast 500 response is not a successful performance result. Performance scripts should validate enough functional behavior to prove that the measured transaction is actually correct.

Separate:

  • protocol failures/timeouts;
  • explicit non-success responses;
  • assertion/check failures where the protocol response is technically successful but business output is wrong.

Under overload, error shape also matters: controlled 429 or queue rejection may be preferable to timeouts, corrupted state or cascading failures.

Resource and saturation metrics

Common server-side signals include:

  • CPU utilization and throttling;
  • memory, heap, garbage collection and allocation rate;
  • thread/worker/connection pools and wait time;
  • queue depth and age;
  • database connections, locks, waits, query latency and I/O;
  • disk latency/IOPS and network saturation;
  • cache hit ratio and eviction;
  • downstream dependency latency/errors;
  • autoscaling and container/VM limits.

Choose from the architecture. A checklist that ignores the actual system is not observability.

Baseline

A baseline records a known system/version under controlled conditions so future runs can be compared. Keep the variables stable enough that a difference is interpretable:

  • same workload model;
  • comparable environment and data;
  • same measurement definitions;
  • known build/configuration;
  • enough repetition to understand normal run-to-run variance.

Thresholds and SLOs

Grafana k6 thresholds turn metrics into pass/fail criteria. A representative example is:

export const options = {
  thresholds: {
    http_req_failed: ['rate<0.01'],
    http_req_duration: ['p(95)<500', 'p(99)<1000'],
  },
};

These numbers are examples of syntax, not universal performance requirements. Replace them with real product objectives.

The value of a threshold is automation: a CI job can fail without a person reading every chart. The limitation is also important: a pass/fail result cannot diagnose the reason for a regression, and a noisy shared environment can create misleading gates. Preserve raw/trended evidence and telemetry.

Correlation is the analysis technique

Suppose p95 rises at minute 8 while throughput plateaus. At the same time:

  • DB connection pool reaches its limit;
  • pool wait time increases;
  • CPU remains moderate;
  • downstream latency is unchanged.

That time-aligned evidence supports a connection-pool/database bottleneck hypothesis. It is far stronger than saying “CPU was high somewhere during the test.”

Reporting rule

Every performance conclusion should name the workload and the relevant metric. “p95 = 420 ms” is incomplete. “At the defined 700 RPS checkout workload, p95 was 420 ms, errors 0.2%, throughput remained stable, and the database pool stayed below saturation” is actionable evidence.

Source registry

Chapter references verified
ISTQB Certified Tester Foundation Level Specialist Syllabus — Performance Testing

Performance-testing terminology, test types, metrics, load generation, operational profiles, planning, execution and analysis

ISTQB · Official syllabus
Source ↗
Grafana k6 — Performance testing fundamentals

Latency, throughput, errors, checks, percentiles and performance-test vocabulary

Grafana Labs · Official documentation
Source ↗
Grafana k6 documentation — Thresholds

Metric pass/fail criteria, percentiles, error-rate limits, SLO encoding and automation

Grafana Labs · Official documentation
Source ↗