4. Testing AI-Based Systems
AI does not remove the need for ordinary software testing. It adds new failure modes and makes some traditional techniques insufficient on their own.
Locked versus adaptive AI systems
A locked model does not change its learned behavior during normal production operation. It may later be retrained and redeployed as a controlled release.
An adaptive system can update behavior after deployment based on new data, feedback, or online learning.
Testing consequence: adaptive behavior increases the importance of continuous monitoring, change detection, rollback, online guardrails, updated acceptance evidence, and controls over what data can influence learning.
Do not assume an externally hosted model is locked merely because your application code is unchanged; the provider may change the model behind an alias unless versioning guarantees say otherwise.
Why statistical testing is necessary
A deterministic function can often be checked input-by-input against exact outputs. ML behavior is evaluated over populations and distributions. You therefore need representative samples and aggregate evidence such as error rates, confidence intervals, per-slice performance, and repeated runs where non-determinism matters.
A single successful AI example is anecdotal. A single failure can still be important, especially for safety or security, but general quality claims require population-level evidence.
The test-oracle problem
A test oracle determines whether observed behavior is acceptable. AI makes oracles difficult when:
- multiple outputs are valid;
- ground truth is expensive or subjective;
- outputs are probabilistic;
- the expected result changes with context;
- a generative output is semantically correct but not textually identical.
Useful alternatives include:
- reference datasets with trusted labels;
- human expert review;
- invariants and business rules;
- metamorphic relations;
- differential/back-to-back comparison;
- statistical thresholds;
- rubrics with multiple quality dimensions;
- consensus or adjudication among reviewers.
An LLM-as-judge can be a tool, but it is not automatically ground truth. Validate the judge, monitor bias/position effects, and use human review for high-risk decisions.
Testing generative AI and LLMs
Separate test dimensions rather than asking “is the chatbot good?”
- Task effectiveness: does it solve the intended task?
- Groundedness/factuality: are claims supported by allowed evidence?
- Relevance: does it address the user request?
- Instruction following: does it respect system/business constraints?
- Safety: does it refuse or safely handle prohibited/high-risk requests?
- Robustness: does behavior survive paraphrases, noise, long context, multilingual input, and adversarial prompts?
- Consistency: how much does output quality vary across repeated samples?
- Security/privacy: prompt injection, data leakage, tool misuse, untrusted content.
- Performance/cost: latency, token use, rate limits, fallback behavior.
Exact-string comparison is appropriate only when the requirement itself is exact, such as a strict JSON enum or mandatory literal token.
Red teaming
Red teaming is adversarial exploration designed to expose harmful, unsafe, insecure or policy-violating behavior. It is broader than ordinary positive/negative functional testing.
A useful red-team campaign varies:
- attacker intent;
- prompt framing and encoding;
- multi-turn escalation;
- indirect prompt injection through retrieved/web content;
- role-play and authority claims;
- tool access and permissions;
- sensitive-data requests;
- language and obfuscation;
- boundary cases between allowed and disallowed behavior.
Record both attack success rate and the quality of safe behavior. A model that refuses every request may be safe against one metric but useless.
Exploratory testing of an LLM
Use a charter instead of random chatting.
Example charter: “Explore whether the customer-support assistant reveals hidden account information when a user changes identity claims across a long conversation.”
Define:
- mission and risk;
- personas and prompt families;
- what evidence counts as a failure;
- variations to try;
- session notes and reproducible transcripts;
- model/configuration/version;
- follow-up automated regression cases for important discoveries.
Hands-on LLM exercise
Pick a public or sandbox LLM and test a narrow requirement: “Answer questions only from the supplied policy text and say when evidence is absent.” Create at least 20 prompts across:
- directly answerable questions;
- paraphrases;
- questions whose answer is absent;
- misleading user assumptions;
- instructions to ignore the policy;
- conflicting text inside the supplied context;
- multilingual variants;
- long-context distractions.
Create a rubric with groundedness, correctness, refusal/abstention correctness, and instruction following. Run important prompts more than once. Summarize failure rate by category rather than reporting one overall “accuracy.”
ML-specific test levels
ML systems introduce test objects that deserve explicit separation. At minimum reason about model-level testing and broader system/integration-level testing alongside conventional component testing.
Model-level evidence asks whether the learned component satisfies functional-performance and robustness expectations on appropriate data. System-level evidence asks whether the entire product uses that model safely and correctly, including preprocessing, postprocessing, fallback, UI/API behavior, permissions, logging and operational conditions.
A model can pass while the system fails—for example, the application swaps class labels or uses the wrong threshold.
Risk-based test strategy for ML
Start with product risks, then map them to lifecycle controls.
Example: resume-screening model
- Risk: discriminatory ranking → representative data checks, subgroup metrics, fairness analysis, human review controls.
- Risk: irrelevant adversarial text manipulates ranking → robustness/adversarial tests.
- Risk: model performance degrades as job market changes → production drift monitoring and periodic labeled evaluation.
- Risk: wrong model version deployed → artifact/version verification and deployment smoke tests.
- Risk: private CV data leaks to third party → privacy/security tests and data-flow review.
The strategy should state scope, risks, test levels, datasets, metrics, thresholds, environments, monitoring, ownership and residual risk—not just a list of test techniques.
Exam traps
- “AI is non-deterministic, therefore exact tests are impossible” is false. Deterministic contracts around the AI still have exact expectations.
- One human reviewer is not a robust oracle for subjective quality.
- Red teaming is not only security penetration testing; it can target harmful or policy-violating model behavior.
- Passing model metrics does not prove system quality.
- Statistical testing does not mean accepting any individual severe failure.
Practice: design a risk-based strategy for an AI meeting-summary product. Include at least five risks, a test level for each, oracle/evaluation method, dataset/sample approach, metric or evidence, and release/monitoring criterion.
Exact ML test-level distinction
CT-AI v2.0 identifies two specialized test levels for ML-specific risks:
- Input data testing — testing the data used for training, testing and prediction, including data quality, representativeness, bias, constraints, labels and pipeline concerns.
- ML model testing — testing the generated ML model itself, including functional performance, robustness and model-specific risks.
Conventional levels still apply where appropriate: component, component integration, system, system integration, and acceptance testing. A likely exam trap is to discard conventional levels because the product contains ML.
Examples:
- test a transformation script in isolation → component testing;
- verify the data pipeline feeds the model the intended feature representation → component integration;
- confirm embedding/compression did not degrade model performance in the complete product → system testing;
- verify exchanges with an external AI service → system integration / relevant API integration testing;
- determine whether a third-party AI service is suitable for the intended product → acceptance testing.
Syllabus hands-on: exploratory LLM + boundary value analysis
In addition to the broader LLM charter earlier in this chapter, perform the syllabus-aligned exercise:
- Give an LLM a clear requirement containing one or more numeric boundaries.
- Ask it to generate test cases using 2-value boundary value analysis.
- Repeat using 3-value boundary value analysis.
- Verify the generated cases manually against the BVA rules.
- Record missing, duplicate, invalid or misclassified boundary cases.
- Change the requirement wording and see whether correctness/completeness changes.
The point is not “use AI to test” as a certification scope. The test object in this exercise is the LLM's ability to perform the requested task correctly and completely.