Locust performance testing — GimmeJob real project walkthrough
This chapter explains Locust through the actual GimmeJob load-test code, not a toy example. The source of truth is `tests/performance/gimmejob/locustfile.py`. The goal is to understand what every important block does, how virtual users create load, how the first Azure run produced about 2.82 requests/second, and how to expand the workload beyond a health check.
1. The mental model
Locust does not open ten browsers. It creates Python virtual users. Each virtual user repeatedly chooses a task, executes an HTTP request, waits according to wait_time, then chooses another task.
create virtual user → on_start() → choose @task → HTTP request → validate response → wait 2–5 s → choose next @task → repeatThe main building blocks are:
class GimmeJobPublicReader(HttpUser):
wait_time = between(2, 5)
@task(4)
def health(self) -> None:
...HttpUserdefines the behavior of one virtual HTTP user.@taskmarks methods Locust may execute repeatedly.- The number in
@task(4)is a weight, not a repeat count. between(2, 5)makes each user wait a random 2–5 seconds between tasks.- The external run configuration controls how many users exist, how quickly they are spawned, how long the run lasts, and which tags are enabled.
2. Imports used by the real script
import logging
import os
from typing import NoReturn
from urllib.parse import urlparse
from locust import HttpUser, between, events, tag, task
from locust.exception import StopUserlogging writes safety and threshold messages. os reads Azure/Locust environment variables. urlparse splits the configured host into scheme and hostname. HttpUser, between, task, tag, and events are the main Locust APIs used by the workload. StopUser terminates a virtual user when a safety guard must stop the run.
3. Safe defaults and production recognition
LOGGER = logging.getLogger(__name__)
LOCAL_HOST = "http://127.0.0.1:4173"
PRODUCTION_HOSTS = {"gimme-job.com", "www.gimme-job.com"}
PRODUCTION_ACKNOWLEDGEMENT = "gimme-job.com"The default target is local development. Production is recognized explicitly, and a production run requires a matching acknowledgement. This is intentional: load generation against production must never happen accidentally.
4. Reading numeric limits from environment variables
The script validates integer configuration with this real helper:
def _positive_int(name: str, default: int) -> int:
raw = os.getenv(name, str(default))
try:
value = int(raw)
except ValueError as error:
raise RuntimeError(f"{name} must be an integer, got {raw!r}.") from error
if value < 1:
raise RuntimeError(f"{name} must be greater than zero.")
return valueEnvironment variables arrive as strings. For example GIMMEJOB_MAX_USERS=10 is read as "10", converted to integer 10, and rejected if it is invalid or less than one.
The float helper follows the same pattern for thresholds such as GIMMEJOB_MAX_FAILURE_RATIO=0.01 and GIMMEJOB_MAX_P95_MS=2500.
5. Reading Locust run configuration
def _configured_user_count(environment: object) -> int:
runner = getattr(environment, "runner", None)
parsed_options = getattr(environment, "parsed_options", None)
values = (
getattr(runner, "target_user_count", 0),
getattr(parsed_options, "num_users", 0),
)
return max(int(value or 0) for value in values)This reads the requested user count from Locust's runtime environment. In the first Azure run the configured value was 10 users.
Duration is read separately:
def _configured_run_seconds(environment: object) -> float | None:
parsed_options = getattr(environment, "parsed_options", None)
value = getattr(parsed_options, "run_time", None)
if isinstance(value, (int, float)):
return float(value)
return NoneFor the first run, LOCUST_RUN_TIME=600 meant 600 seconds, or 10 minutes.
6. Emergency stop
def _stop_run(user: HttpUser, reason: str) -> NoReturn:
LOGGER.error("GimmeJob load-test safety guard stopped the run: %s", reason)
runner = getattr(user.environment, "runner", None)
if runner is not None:
runner.quit()
raise StopUser()If a production guard fails, runner.quit() stops the overall run and StopUser terminates the current virtual user.
7. One class equals one virtual-user behavior
class GimmeJobPublicReader(HttpUser):
"""A realistic public visitor that never sends a mutating request."""
host = os.getenv("GIMMEJOB_HOST", LOCAL_HOST).rstrip("/")
wait_time = between(2, 5)When Azure is configured for 10 users, Locust creates ten instances of this behavior. They are lightweight Python virtual users, not browser sessions.
host resolves to https://gimme-job.com in the production run. Therefore:
self.client.get("/api/health")actually sends:
GET https://gimme-job.com/api/health8. Why the first run produced about 2.82 RPS
wait_time = between(2, 5)The average wait is about 3.5 seconds. With 10 users, a rough steady-state estimate is:
10 users / 3.5 seconds ≈ 2.86 requests/secondAzure reported about 2.82 requests/second, almost exactly what the workload model predicts. The 2.82 RPS was therefore not the site's maximum capacity; it was the rate intentionally produced by the current virtual-user pacing.
9. on_start() runs once per virtual user
def on_start(self) -> None:
parsed_host = urlparse(self.host)
hostname = (parsed_host.hostname or "").lower()
if hostname in PRODUCTION_HOSTS:
if parsed_host.scheme != "https":
_stop_run(self, "production tests require an https:// host")
acknowledgement = os.getenv("GIMMEJOB_PRODUCTION_ACK", "")
if acknowledgement != PRODUCTION_ACKNOWLEDGEMENT:
_stop_run(
self,
"set GIMMEJOB_PRODUCTION_ACK=gimme-job.com before targeting production",
)For each new virtual user, the script parses the target and checks that a production run uses HTTPS and has the explicit production acknowledgement.
It also checks that requested users and duration stay within configured limits:
max_users = _positive_int("GIMMEJOB_MAX_USERS", 10)
configured_users = _configured_user_count(self.environment)
if configured_users > max_users:
_stop_run(self, ...)
max_run_seconds = _positive_int("GIMMEJOB_MAX_RUN_SECONDS", 600)
configured_run_seconds = _configured_run_seconds(self.environment)
if configured_run_seconds is None:
_stop_run(self, ...)
if configured_run_seconds > max_run_seconds:
_stop_run(self, ...)The first production configuration was 10 users and 600 seconds, so both guards passed exactly at their limits.
Finally, each user's HTTP client receives identifying headers:
self.client.headers.update(
{
"accept": "application/json, text/html;q=0.9",
"user-agent": "GimmeJob-authorized-Locust/1.0",
}
)The dedicated User-Agent helps distinguish authorized Locust traffic in server or Cloudflare logs.
10. Locust can decide that HTTP 200 is still a failure
The script uses catch_response=True:
with self.client.get(path, name=name, catch_response=True) as response:
if response.status_code != 200:
response.failure(_status_failure(response))
returnThis allows application-level validation. A response may be 200 OK but still be wrong—for example, invalid JSON, the wrong HTML page, or a missing contract field. Calling response.failure(...) records that request as failed in Locust statistics.
11. The exact health task used in the first Azure run
@tag("health", "smoke", "api", "worker")
@task(4)
def health(self) -> None:
with self.client.get("/api/health", name="GET /api/health", catch_response=True) as response:
if response.status_code != 200:
response.failure(_status_failure(response))
return
try:
payload = response.json()
except ValueError:
response.failure("health response is not valid JSON")
return
if (
not isinstance(payload, dict)
or payload.get("ok") is not True
or payload.get("service") != "jobpilot-cloud"
):
response.failure("health response does not match the public contract")The request is successful only if all of these are true:
- HTTP status is 200.
- The body is valid JSON.
- The JSON root is an object.
okis exactlytrue.serviceequalsjobpilot-cloud.
Therefore the first run's 0 errors meant more than “the server responded”: all completed responses matched this health contract.
12. Tags select scenarios
The first Azure run used LOCUST_TAGS=smoke. Because the health task has the smoke tag, only that task was eligible. This is why Azure statistics showed only GET /api/health.
The same file already contains other scenarios.
Public home page
@tag("home", "public-read", "edge", "html")
@task(3)
def home_page(self) -> None:
self._expect_html("/", "Why I created this site", "GET / [public home]")This exercises normal public page delivery and edge caching. It validates HTTP 200, text/html, and a known page marker.
Uncached reference page
@tag("reference", "public-read", "worker", "html")
@task(2)
def uncached_reference_page(self) -> None:
self._expect_html(
"/reference/qa-fundamentals",
"Core QA distinctions",
"GET /reference/qa-fundamentals [uncached]",
)This exercises Worker-rendered HTML rather than the lightweight health endpoint.
Public jobs — D1 read path
@tag("jobs", "public-read", "d1", "api")
@task(2)
def public_jobs(self) -> None:
with self.client.get(
"/api/public/jobs",
name="GET /api/public/jobs [D1]",
catch_response=True,
) as response:
...This exercises the public D1 database read path and validates that the response contains a jobs list and generatedAt.
Dashboard — heavier D1 read path
@tag("dashboard", "public-read", "d1", "api", "heavy")
@task(1)
def public_dashboard(self) -> None:
with self.client.get(
"/api/dashboard",
name="GET /api/dashboard [D1 heavy]",
catch_response=True,
) as response:
...This is the heavier public D1 scenario and validates jobs plus market.
13. Task weights model traffic mix
The real task weights are:
| Task | Weight | Approximate share when all tasks are enabled |
|---|---|---|
| Health | 4 | 33.3% |
| Home page | 3 | 25.0% |
| Reference page | 2 | 16.7% |
| Public jobs | 2 | 16.7% |
| Dashboard | 1 | 8.3% |
Weights only matter when several tasks are eligible. With LOCUST_TAGS=smoke, health is the only eligible task, so its weight does not change anything.
14. Spawn rate is not requests per second
For the first run:
users = 10
spawn rate = 1
duration = 600 sspawn rate = 1 means create one new virtual user per second until the target of 10 users is reached. It does not mean 1 request/second. RPS emerges from the number of active users, task execution time, and wait_time.
15. End-of-run thresholds
The script attaches a listener to Locust's quitting event:
@events.quitting.add_listener
def apply_exploratory_thresholds(environment: object, **_kwargs: object) -> None:It reads total statistics and calculates p95:
max_failure_ratio = _non_negative_float("GIMMEJOB_MAX_FAILURE_RATIO", 0.01)
max_p95_ms = _non_negative_float("GIMMEJOB_MAX_P95_MS", 2500)
p95_ms = stats.get_response_time_percentile(0.95) or 0The default contract is therefore:
- failure ratio <= 1%;
- aggregate p95 <= 2500 ms.
If a threshold is exceeded, the process exit code becomes 1 so Azure/CI can treat the run as failed.
16. Reconstructing the first run
The first Azure configuration was effectively:
Host: https://gimme-job.com
Users: 10
Spawn rate: 1 user/s
Duration: 600 s
Tags: smoke
Wait time: random 2–5 s per userExecution looked approximately like this:
0 s user #1 → on_start → health loop
1 s user #2 → on_start → health loop
2 s user #3 → on_start → health loop
...
9 s user #10 → on_start → health loop
all active users → GET /api/health → validate → wait 2–5 s → repeatThe observed result—1701 requests, 0 errors, p90 about 21 ms, throughput about 2.82 RPS—is consistent with this workload model. It is a successful smoke/load-tool validation, not a capacity limit for GimmeJob.
Where the test actually runs
The Python file in GitHub defines the workload, but GitHub is not the place where the production load is generated. In the Azure setup, Azure Load Testing provisions a test engine and starts Locust on that engine. Locust then sends real HTTPS requests to https://gimme-job.com.
Azure Load Testing engine → Locust virtual users → https://gimme-job.com → Cloudflare edge → Worker gimmejob → D1 gimmejob-db when needed → response → Azure/Locust statisticsThis separates the load generator from the system under test:
- Azure Load Testing runs Locust and creates the virtual users.
gimme-job.comis the target being measured.- Cloudflare serves the requests and exposes server-side/runtime telemetry.
- The Worker
gimmejobhandles the application request. GET /api/public/jobsandGET /api/dashboardalso exercise D1gimmejob-db.- A local Locust run changes only the load-generator location: Locust runs on the developer machine instead of an Azure engine.
The first Azure run used LOCUST_TAGS=smoke, so only the health task was eligible. The current workload also has a public-read selector. Setting LOCUST_TAGS=public-read makes Locust select among the four implemented non-health reads: GET /, GET /reference/qa-fundamentals, GET /api/public/jobs, and GET /api/dashboard. Route-specific selectors home, reference, jobs, and dashboard allow one path to be isolated.
Where to read performance results
A useful performance test has two views of the same run. Azure/Locust shows what the load generator experienced; Cloudflare shows what happened inside the platform serving those requests.
| Question | Primary place to look |
|---|---|
| How many requests were sent and completed? | Azure Load Testing run dashboard |
| What were p50/p95/p99 response times? | Azure Load Testing statistics |
| What throughput/RPS did the test actually achieve? | Azure Load Testing |
| Which request failed or returned the wrong contract? | Azure/Locust errors and named-request statistics |
| Did the Worker hit runtime or CPU problems? | Cloudflare Workers metrics |
| Were there Worker exceptions or resource errors? | Cloudflare Workers metrics/logs |
| Did D1 queries become slower or scan more rows? | Cloudflare D1 analytics |
| What probably caused a latency increase? | Correlate Azure and Cloudflare over the same timestamps |
For GimmeJob, open the Azure Load Testing resource, select the test, then the concrete test run. That run is the source for client-side response time, throughput, request count and failure statistics.
In Cloudflare, use Workers & Pages → gimmejob for Worker metrics/Observability. Inspect request counts, invocation statuses/errors, CPU time, wall/execution time and logs/traces when needed. For D1-backed routes, also inspect the gimmejob-db analytics for read query rate, rows read and query latency.
The first health-only run can be understood mainly from Azure because its purpose was to validate the script and generated traffic. Once D1/public routes are exercised, Azure alone is not enough to diagnose a slowdown. Use Azure and Cloudflare together and align their graphs by the same test start/end timestamps.
Official references: Azure Load Testing, Cloudflare Workers metrics, and Cloudflare D1 metrics.
Summary
- A Locust
HttpUseris a lightweight virtual user, not a browser. @taskdefines repeatable behavior; task weights define relative selection probability.- Tags allow one scenario or a group of scenarios to be isolated.
wait_timeis a major part of the workload model and directly affects RPS.- Spawn rate controls how quickly users appear, not request rate.
catch_response=Trueallows contract failures even when HTTP status is 200.- The first GimmeJob run tested only the
smokehealth task becauseLOCUST_TAGS=smokewas enabled. - The existing script already includes edge HTML, uncached Worker HTML, D1 jobs, and heavier D1 dashboard scenarios.
- Thresholds convert performance expectations into a pass/fail test contract.