Bottlenecks, scalability and recovery
A performance test becomes engineering evidence when it can explain the relationship between workload, user-visible behavior and a constrained resource.
First rule: validate the measurement system
Before blaming the application:
- confirm the generator achieved the intended load;
- confirm generator CPU/memory/network have headroom;
- confirm scripts are succeeding functionally;
- confirm data collisions or rate limits are not producing artificial errors;
- confirm the environment and build are the ones you intended to measure.
Only then interpret target-system limits.
Find the inflection point
Increase load in controlled stages and look for a change in relationship:
- latency rises sharply;
- error rate increases;
- useful throughput plateaus or falls;
- queue depth begins to grow without recovering;
- one resource approaches sustained saturation.
The load level and timestamp of that transition are the anchor for diagnosis.
Correlate across layers
At the same point, inspect likely constraints:
Application/runtime
- CPU throttling or saturation;
- allocation/GC pauses;
- worker or thread pool exhaustion;
- lock contention;
- event-loop delay;
- internal queue growth.
Database
- connection-pool waits;
- slow or changing query plans;
- locks/deadlocks/contention;
- cache/buffer behavior;
- I/O latency;
- replication or commit latency.
Infrastructure/network
- container/VM CPU limits;
- memory pressure/OOM;
- disk I/O;
- network bandwidth/connections;
- load balancer limits;
- autoscaling delays.
Dependencies
- external API latency/errors;
- queue or broker lag;
- downstream rate limits;
- retry amplification.
Prove causality by controlled change
A correlation creates a hypothesis. Change one credible constraint and rerun the same workload. Examples:
- increase a connection pool only after proving pool wait;
- add an index only after showing the relevant query/plan;
- increase CPU only after proving CPU saturation;
- disable or tune a retry only after showing amplification.
If the limit moves in the predicted direction, the evidence is stronger.
Scalability
Measure useful capacity as resources are added. Important questions:
- Does throughput increase proportionally?
- Does latency stay inside the objective?
- Does a shared database or dependency become the new bottleneck?
- Does autoscaling react before queues become unsafe?
- What is the cost per useful transaction at each scale?
Horizontal scale can expose shared-state, locking, partitioning or coordination costs that a single-node test never sees.
Graceful degradation
A system does not need to remain fully functional at unlimited load. It should fail in a controlled way consistent with design. Depending on the product, acceptable overload mechanisms may include:
- bounded queues;
- explicit
429/backpressure; - load shedding;
- degraded non-critical features;
- circuit breakers;
- cached/static fallback.
The performance test should verify correctness during degradation. “It stayed up” is insufficient if it accepted requests and lost data.
Recovery
After a spike or stress phase, verify the system returns to a healthy steady state:
- queues drain;
- error rate returns to normal;
- latency distribution returns to baseline range;
- pools/resources are released;
- autoscaled capacity settles as designed;
- no retry storm continues;
- data remains consistent.
Recovery time can itself be a requirement.
Endurance diagnosis
For long tests, trend values rather than inspecting only start/end snapshots. A memory leak may appear as a repeating sawtooth with a rising floor; a connection leak as a pool's available count that never fully recovers; a queue leak as depth whose minima rise each cycle.
Diagnosis output
A strong report separates fact from hypothesis:
Observed: At 700–800 RPS, p95 increased from the steady range to 1.8 s, useful throughput stopped increasing, and DB pool wait time rose while all pool connections were occupied.
Hypothesis: Database connection capacity or downstream DB work is limiting throughput.
Next experiment: Capture slow-query/DB-wait evidence and rerun the same workload after one controlled pool/query change.
That is more useful than “the application is slow under load.”