Foundations & choosing a layer
What automation is for, the layers available, and deciding what to automate at all.
What test automation is actually for
Automation buys fast, repeatable feedback — not thoroughness, and not a replacement for testers.
An automated test is a check: it compares the system against an expectation somebody already had. That makes automation excellent at one specific job — telling you quickly and repeatedly whether behaviour that used to work still works. It is structurally incapable of noticing anything nobody predicted, because an assertion can only fail against an expectation that was written down. Teams that treat a suite as proof of quality end up with hundreds of green tests and users reporting problems nobody thought to assert on. The honest framing is narrower and more useful: automation protects known behaviour so that humans have time to investigate the unknown. Every test therefore needs an answer to one question — what risk does this cover? If the answer is "none, but it adds coverage", the test is pure maintenance cost.
Key points
- Automation checks known expectations; it cannot be surprised by something nobody wrote down
- Its real product is fast, repeatable feedback, which is what makes frequent releases safe
- Every test should name the risk it covers, or it is maintenance cost without a benefit
Code
# Two tests, same endpoint. Only one of them earns its maintenance cost.
def test_cart_total_is_sum_of_line_totals(shop_as_user):
"""Risk: wrong money shown to a customer. Worth automating forever."""
products = shop_as_user.products.list(in_stock=True)[:3]
for index, product in enumerate(products, start=1):
shop_as_user.cart.add_item(product.id, index)
cart = shop_as_user.cart.get()
expected = sum(p.price * i for i, p in enumerate(products, start=1))
assert cart.total == pytest.approx(expected, abs=0.01)
def test_cart_response_contains_a_currency_field(shop_as_user):
"""Risk: none anybody can name. Deleting this test loses nothing."""
assert "currency" in shop_as_user.cart.get().model_dump()The first test protects a real risk; the second exists only to raise a coverage number.
See it in the framework
Reviewed against commit f9ada16 (2026-08-14).
Common pitfalls
- Treating a green suite as evidence of quality rather than evidence that known behaviour is intact
- Automating test cases one-for-one from a manual regression sheet without asking which risks still matter
Practice exercise
Take any five tests from a suite you know and write down, in one sentence each, the risk they cover. Any test where you cannot name a risk is a candidate for deletion — decide honestly whether you would still write it today.
Choosing the layer: services, UI, or neither
Push every test down until it can no longer prove what you need — the cost difference between layers is enormous.
The same business rule can usually be checked at several levels, and the cost difference is not marginal. An API test that verifies a cart total runs in about a hundred milliseconds, has no rendering to go wrong, and fails with the exact request and response attached. The equivalent browser test takes seconds, depends on layout, timing and network, and fails with a screenshot you still have to interpret. Neither is better in the abstract; what matters is asking which is the cheapest layer that can still prove the thing you care about. Business rules, validation, authorization and error handling belong at the service layer. The user interface layer should be reserved for what only it can prove: that the interface sends what the backend expects, renders what it receives, and remains usable. Teams that invert this end up with a ninety-minute pipeline and dozens of flaky failures a week, and every one of those tests had a cheaper equivalent one layer down.
Key points
- Ask of every test: what is the cheapest layer that can still prove this?
- Business rules, boundaries and authorization belong to the service layer
- Reserve UI tests for wiring, rendering and usability — what only the interface can prove
Code
# The rule (a quantity above 99 is rejected) belongs here — 100 ms, no browser.
@pytest.mark.parametrize(
("quantity", "expected_status"),
[(1, 201), (99, 201), (0, 422), (100, 422), (1.5, 422)],
)
def test_quantity_boundaries(shop_as_user, quantity, expected_status):
product = shop_as_user.products.first_in_stock()
response = shop_as_user.cart.add_item_raw(product.id, quantity)
assert response.status_code == expected_status
# The UI test that remains proves only what the API test cannot: that the
# interface actually shows the rejection to a human being.
def test_rejected_quantity_is_shown_to_the_user(app_as_user):
products = app_as_user.products.open()
products.cards[0].add_to_cart()
expect(products.status).to_contain_text("added to cart")Five boundary cases at the API layer, one wiring check in the browser — not five browser tests.
See it in the framework
- Used by · Service-layer test:
tests/services/test_cart_flow.py— live · reviewed - Used by · Same flow at the UI layer:
tests/web/test_shopping_flow.py— live · reviewed
Reviewed against commit f9ada16 (2026-08-14).
Common pitfalls
- Automating at the UI layer because it maps neatly onto an existing manual test case
- Adding a second end-to-end journey when the first already covers the wiring, so the new one only adds runtime and flake
Practice exercise
Find a UI test in a suite you work with that asserts a calculation or a validation rule. Rewrite it as an API test, then decide what — if anything — the UI test should still assert once the rule is covered below.
Deciding what to automate — and what to leave alone
Score candidates on risk, frequency, repetition, determinism and stability; be willing to say no.
A candidate for automation scores well when the risk is high, the check runs often, the steps repeat with different data, the outcome is deterministic, and the feature has stopped changing shape every sprint. It scores badly when any one of those is missing, and a low score on stability alone is usually enough to wait — automating an interface that is still being designed means writing the test three times. A few things resist blanket automation, though the reason differs case by case: whether a layout looks right, or whether error copy is clear, is a judgement call a human has to make at least once — but the resulting bar can often be captured afterward as an automatable check (a visual-regression diff against an approved baseline, a length or tone lint on copy) rather than staying manual forever. A one-off migration is rarely worth a bespoke test suite for its own sake, but the reconciliation it produces — row counts, checksums, before/after diffs — is exactly the kind of check automation is good at. Exploratory testing is the one genuinely irreducible case: it depends on a human being able to be surprised, which a fixed script cannot do. Saying "we should not automate that" is a senior skill, and it is easier to defend when you can point at the scoring. The cost people forget is not writing the test — it is every future engineer who has to read, debug and update it.
Key points
- Score candidates on risk, frequency, repetition, determinism, manual cost and stability
- Do not automate an interface that is still being redesigned — wait for the second iteration
- Exploratory testing cannot be automated, because automation cannot be surprised
Code
Run the code to see the result.
Risk equals probability multiplied by impact; the row that says "None" is the one that proves you prioritised.
See it in the framework
- Strategy notes:
docs/09-strategy.md— live · reviewed - Used by · Visual/accessibility checks:
tests/web/test_accessibility_and_visual.py— live · reviewed
Reviewed against commit f9ada16 (2026-08-14).
Common pitfalls
- Chasing an automation percentage target, which says nothing about whether the uncovered part holds the risk
- Automating a brand-new screen during the sprint it is being built, then rewriting the test twice
Practice exercise
Build this table for a product you know: every significant feature, scored for probability and impact, with the layer you would cover it at. Include at least one row where the answer is deliberately no automation, and be ready to defend it.