7. ML Development & Deployment Testing
The shortest chapter covers a failure mode that causes expensive incidents: the model that was evaluated is not necessarily the model that is packaged, configured and serving production traffic.
ML development risks
Treat the ML build as a versioned system of artifacts:
- source code;
- data snapshot/query and labels;
- preprocessing/feature code;
- training configuration and random seeds;
- framework/library versions;
- model architecture/configuration;
- learned weights/model artifact;
- evaluation dataset and results;
- conversion/quantization step;
- serving container/runtime;
- application thresholds and postprocessing.
Risks include non-reproducible builds, dependency drift, wrong data versions, accidental use of test data for tuning, untracked feature changes, artifact corruption, wrong model/configuration deployed, and train/serve skew.
A robust pipeline records enough lineage to answer: which code + data + configuration produced this exact model, which tests passed, and what is serving now?
Deployment testing
Deployment tests should prove more than “the endpoint returns 200.”
Check:
- exact model/artifact/version loaded;
- preprocessing and feature schema match training expectations;
- postprocessing/threshold/configuration match approved values;
- representative smoke inputs produce plausible expected behavior;
- latency, memory and hardware behavior on the target runtime;
- permissions, secrets, network dependencies and external model endpoints;
- logging/telemetry include model/version identifiers and useful diagnostics without leaking sensitive data;
- rollback/fallback path works;
- canary or staged rollout routes traffic correctly;
- monitoring for data/model/system health is live before full exposure.
Example deployment defect: model B passes offline evaluation, but production loads model A with model B’s threshold. Both artifacts are individually valid; the deployed combination is not.
Minimal deployment evidence record
release: 2026.08.26
model_sha256: ...
model_version: fraud-v18
training_data_version: transactions-2026-07-31
preprocessing_commit: abc1234
runtime_image: fraud-serving@sha256:...
threshold: 0.73
acceptance_dataset: fraud-acceptance-v7
acceptance_result: PASS
canary_result: PASS
rollback_target: fraud-v17The exact tooling is not important for the exam. The principle is traceability and testability across the ML lifecycle.
Practice: write a deployment test for a model served behind an API. Include artifact identity, schema, five fixed smoke cases, latency, monitoring, canary traffic and rollback. Then state which failures should block rollout immediately.
Deployment-testing taxonomy to memorize
For AI-7.1.2, recognize the named deployment test types and the risk each addresses:
- Installability testing: installation, configuration, dependencies and supported environments.
- Rollback testing: prove the model/system can return to a known stable version after a bad rollout.
- Canary testing: expose a small controlled portion of real traffic to the new deployment and watch agreed metrics before expansion.
- Shadow testing: send live requests to the new model in parallel without letting its result affect the live response; compare it with the current model.
- Model conversion testing: after converting/compressing/quantizing to a deployment format, re-check predictive behavior and operational efficiency.
- Cross-device / device-compatibility testing: verify intended devices/edge/cloud targets behave acceptably.
- API testing: verify the deployed ML service contract, input/output handling, errors and integration behavior.
Canary vs shadow vs A/B vs back-to-back
- Canary: small live exposure to the new deployment; it can affect those selected users.
- Shadow: duplicate live traffic to the candidate model, but its answer does not control the user-facing result.
- A/B: controlled live experiment comparing alternatives as treatments and outcomes.
- Back-to-back: feed comparable/same inputs to implementations and analyze output differences to detect defects/regressions.
These techniques can be combined in a rollout, but they answer different questions.