Workload modelling
A load tool can generate traffic very precisely and still produce a meaningless test if the traffic model is wrong. Workload modelling is the step where production behavior becomes an executable test.
Start from an operational profile
Use evidence where possible:
- production request or transaction rates;
- analytics for user journeys and session duration;
- endpoint or business-operation mix;
- expected growth and peak multipliers;
- regional/time-of-day patterns;
- read/write ratios;
- data-size distribution;
- cache hit/miss behavior;
- background jobs and asynchronous traffic.
If production data does not exist, document the business assumptions. An assumption is acceptable when it is visible; an unexplained round number is not.
Concurrency is not throughput
Concurrent users is a population: how many users are active at the same time. Throughput is a rate: how many requests or business transactions complete per second or minute. A “hit” or request is an event, not a rate.
The same 100 users can generate radically different throughput depending on response time, think time, pacing, parallel requests and the journey itself. This is why every run should record the traffic that was actually achieved, not just the configured VU count.
Think time and pacing
Real people do not usually execute requests in a zero-delay loop.
- Think time models delay between user actions.
- Pacing controls how frequently a full iteration or business flow repeats.
Removing them can be correct for a component capacity test, but then the test should be described as a capacity-oriented protocol test rather than a realistic user model.
Closed workload model
In a closed model, a virtual user starts its next iteration only after the previous iteration completes. If the system becomes slower, each VU completes fewer iterations, so arrival rate naturally falls.
This is appropriate when the real population is bounded and users wait for completion before doing more work. It can be misleading when real arrivals continue independently of response time.
Open workload model
In an open model, new iterations arrive according to an external rate. The start rate is decoupled from how long earlier iterations take. This is useful for traffic such as incoming API calls, messages or users arriving at a service regardless of current latency.
Grafana k6 exposes this distinction directly: VU executors such as constant-vus are closed, while arrival-rate executors such as constant-arrival-rate and ramping-arrival-rate are open.
Coordinated omission
If a closed generator waits for a slow system before sending the next iteration, the system's slowdown also reduces offered load. That can under-sample the period when the service is struggling and make results look better than an arrival process that continues at the expected rate. This class of measurement error is commonly called coordinated omission.
The fix is not “always use open model.” The fix is to choose the model that matches the real workload and understand the measurement consequence.
Load shape
A typical run has deliberate phases:
- warm-up when caches, pools, JIT or runtime state may settle;
- ramp-up to avoid confusing startup transients with steady-state behavior;
- steady state long enough to collect a useful distribution;
- ramp-down/recovery where relevant.
A spike test intentionally changes this shape. A stress test may add progressively higher plateaus. An endurance test extends the steady period.
Traffic mix
Suppose production peak traffic is:
- browse: 60%;
- search: 25%;
- checkout: 15%.
A realistic model should preserve those proportions at the relevant rate. Making every user execute browse → search → checkout in a fixed loop may accidentally create a 33/33/33 mix and overstate writes.
Locust example: weighted behavior
from locust import HttpUser, between, task
class Shopper(HttpUser):
wait_time = between(1, 4)
@task(4)
def browse(self):
self.client.get("/api/products")
@task(2)
def search(self):
self.client.get("/api/search", params={"q": "laptop"})
@task(1)
def cart(self):
self.client.get("/api/cart")The weights describe relative task selection, and wait_time adds user delay. This is only a model; use real traffic evidence to choose the weights and timing.
A headless run can then explicitly state the population and spawn rate:
locust -f locustfile.py --headless -H https://example.test -u 100 -r 10 -t 10mHere -u is peak concurrent Locust users, -r is users spawned per second, and -t is run time. None of those parameters tells you the achieved RPS; measure it.
Workload-model checklist
Before execution, be able to answer: What business operations are represented? In what proportions? At what arrival/transaction rate or concurrency? With what timing and data? For how long? What production evidence supports the model? What actual throughput must the generator demonstrate during the run?