CI/CD, regression and reporting
Performance testing belongs in the delivery lifecycle, but not every performance test belongs on every pull request. The solution is a layered strategy.
Layer 1: cheap performance smoke
Purpose: catch obvious regressions quickly.
Characteristics:
- short duration;
- small controlled load;
- critical endpoint or journey only;
- stable isolated environment if possible;
- explicit correctness and performance thresholds.
This test does not establish production capacity. It is a regression signal.
Layer 2: scheduled representative load
Run a production-representative traffic mix at meaningful load on a controlled environment. Nightly or scheduled execution can be appropriate when the environment and runtime cost are too high for every commit.
Compare against:
- product/SLO objectives;
- a known-good baseline;
- recent trend distribution.
Investigate unexplained variance rather than simply widening thresholds until the pipeline stops failing.
Layer 3: risk-driven large tests
Stress, spike, endurance, large distributed and major capacity tests are usually event/schedule driven:
- before a major traffic event;
- after architecture or infrastructure changes;
- after database/data-volume changes;
- before a release with performance risk;
- periodically to validate capacity assumptions.
Thresholds as gates
Automated thresholds should state both the metric and the workload they apply to. For example, a p95 target from a 20-VU smoke cannot be interpreted as the production peak SLO unless the two workloads are intentionally equivalent.
k6 makes gates explicit in the script. JMeter can be integrated into pipelines by processing JTL/HTML/backend metrics and applying project-defined criteria. Whichever tool is used, the requirement is the same: a machine-readable decision plus retained evidence for diagnosis.
Reproducibility
For a regression comparison, record:
- build/commit/version;
- test-script version;
- environment and configuration;
- workload parameters;
- test data revision;
- start/end time;
- generator topology;
- relevant service/infrastructure changes.
Without this metadata, a faster or slower run may simply be a different experiment.
Repeatability and noise
Shared cloud environments, autoscaling, neighbors, scheduled jobs and external dependencies create variance. For important regression gates:
- understand normal run-to-run variance;
- repeat suspicious results;
- compare equivalent load phases;
- avoid mixing warm and cold states unintentionally;
- keep the environment as deterministic as practical;
- do not hide instability by reporting only averages.
Report for decisions, not decoration
A compact performance report should answer:
- What was tested? Build, environment, architecture scope.
- Under what workload? Traffic mix, rate/concurrency, timing, duration, data.
- What were the objectives? SLO/threshold/baseline and why those numbers exist.
- What happened? Percentile latency, throughput, errors, correctness.
- What constrained the system? Correlated resource/dependency evidence.
- Did it recover? When the test includes overload.
- What changed from baseline? Comparable result, not an unrelated run.
- What decision follows? Pass, investigate, tune, scale, defer with risk, or rerun.
Example result statement
Weak: “Average response time was 320 ms and the test passed.”
Useful: “Under the defined checkout peak workload, achieved traffic remained at the target rate; p95 and error rate stayed within the agreed objectives through the steady phase. p99 worsened relative to the last known-good build while DB pool wait increased at the same timestamps. The release gate passed, but the tail-latency regression is retained for investigation before capacity testing.”
Final workflow
A mature performance process is iterative:
requirements/risk → workload model → validated scripts/data → controlled execution → client and server metrics → diagnosis → engineering change → same-workload rerun → baseline/trend → delivery decision.
Tools automate the middle. The quality of the conclusion still depends on the model, measurement boundaries and evidence.