GimmeJob
Sign in
Certification learning path · CT-AI v2.0 · Chapter 07 / 12

ISTQB CT-AI v2.0 Exam Preparation

6. Model Testing

Model testing asks whether the learned component behaves acceptably before and after integration. It should be traceable to risks, data assumptions and measurable acceptance criteria.

Model risks and mitigations

Typical risks include:

  • poor functional performance on important classes/slices;
  • instability near decision boundaries;
  • sensitivity to irrelevant changes/noise;
  • adversarial manipulation;
  • overfitting and weak generalization;
  • underfitting and insufficient capacity/features;
  • drift after deployment;
  • inappropriate confidence/calibration;
  • undocumented data/usage limitations;
  • reproducibility/versioning failures.

Mitigations combine better data/model development with independent testing, robust acceptance criteria, monitoring and safe system controls.

Model documentation and review

Before executing black-box tests, review evidence. Useful artifacts include:

  • intended use and prohibited/out-of-scope use;
  • training and evaluation data description;
  • model version and architecture/family;
  • preprocessing and feature definitions;
  • metrics overall and by important slice;
  • known limitations;
  • threshold choices;
  • validation procedure;
  • reproducibility information;
  • safety/security considerations;
  • monitoring and retraining triggers.

A review can find testability gaps early: no frozen acceptance dataset, no model version ID, missing subgroup metrics, or an undocumented threshold.

Functional performance of probabilistic models

Do not judge a probabilistic model using a few exact cases alone. Evaluate over a dataset that matches the intended population and calculate suitable metrics.

Questions to ask:

  • Is the acceptance dataset independent of training/tuning?
  • Is it representative and sufficiently large?
  • Are rare/high-risk classes analyzed separately?
  • Are thresholds fixed before final evaluation?
  • Are confidence intervals or repeated samples needed?
  • Do metrics match the product’s error costs?

For generative models, use repeated samples and multidimensional rubrics where output variability is relevant.

Adversarial testing

Adversarial testing deliberately searches for inputs that cause incorrect or unsafe behavior. Depending on the system, attacks can be tiny image perturbations, crafted feature values, prompt injection, malicious retrieved documents, evasion patterns, poisoning attempts, or sequences of actions.

The purpose is not merely to demonstrate that an attack exists. Characterize preconditions, impact, success rate, detectability, mitigations and residual risk.

Metamorphic testing

Metamorphic testing is powerful when exact expected outputs are unavailable. Instead of asserting one exact answer, define a relation between outputs for related inputs.

Examples:

  • Slightly increase image brightness within the supported range → traffic-sign class should remain unchanged.
  • Reorder independent items in a set → aggregate prediction should remain equivalent where order has no semantics.
  • Translate a simple supported-language intent → semantic classification should remain the same.
  • Add irrelevant whitespace to a structured text field → result should not materially change.
  • Increase a clearly risk-increasing feature while holding everything else constant → risk score should not decrease, if that monotonic relation is a valid domain requirement.

Metamorphic relations must come from real requirements/domain properties. Do not invent invariants the model was never intended to satisfy.

Hands-on metamorphic test

Suppose a model exposes:

Python
editable · browser sandbox
Result
Run the code to see the result.

Create transformations for brightness +5%, JPEG recompression, one-pixel translation, and harmless metadata removal. For each transformation define:

  • applicability precondition;
  • expected relation (same class, confidence delta within tolerance, etc.);
  • number/type of seed images;
  • failure evidence;
  • whether one failure is critical or should be evaluated statistically.

Drift

Data drift means the input distribution changes. Concept drift means the relationship between inputs and the target changes. Model performance may degrade even if code and model files are unchanged.

Monitor leading indicators (input distributions, missingness, category changes) and outcome indicators (performance on newly labeled production data). Define thresholds that trigger investigation, re-evaluation or retraining.

Drift detection does not automatically prove performance degradation; it tells you an assumption changed and evidence should be refreshed.

Overfitting and underfitting

Overfitting: excellent training performance but weaker performance on unseen data. The model learned training-specific patterns/noise.

Underfitting: poor performance even on training data; the model/features/capacity/training are insufficient for the problem.

A common diagnostic pattern:

training high, validation much lower -> suspect overfitting
training low, validation similarly low -> suspect underfitting

Do not diagnose from one number alone; verify dataset quality, split strategy, leakage and metric choice.

A/B testing

A/B testing compares alternatives with different users/traffic under controlled assignment. It is useful when offline metrics do not fully predict product outcomes.

Key test concerns:

  • randomization/assignment is correct;
  • populations are comparable;
  • experiment duration/sample size is adequate;
  • primary and guardrail metrics are pre-defined;
  • exposure does not create unsafe treatment;
  • results are not cherry-picked after many metrics/segments are inspected.

A/B is an online experiment, not merely “run two models on the same file.”

Back-to-back testing

Back-to-back (differential) testing sends the same or equivalent inputs to two implementations/models and compares outputs. One may be a previous model, reference implementation, simpler model or alternate provider.

This is especially useful when exact oracles are difficult. Differences reveal where investigation is needed, but disagreement alone does not tell you which model is correct.

Example: before replacing model A with model B, run both across a frozen regression corpus. Compare class changes, confidence shifts and slice-level metrics. Investigate large regressions even if B’s global accuracy is higher.

Exam traps

  • Metamorphic testing checks relations, not exact known outputs.
  • A/B testing uses separate live treatments; back-to-back uses comparable inputs for direct output comparison.
  • Drift is not necessarily a code defect.
  • Overfitting is not the same as high model complexity in isolation.
  • Adversarial tests need threat assumptions and impact, not random malformed inputs only.

Practice: for a new version of an image moderation model, write one test each for functional performance, adversarial behavior, metamorphic behavior, drift readiness, overfitting evidence, A/B readiness and back-to-back regression. State the oracle for every test.

Complete model-risk test toolbox

When AI-6.1.1 asks for test approaches that mitigate model risks, be ready to recognize the broader toolbox, not only metric testing:

  • testing for bias / ethical-system concerns;
  • adversarial testing;
  • overfitting and underfitting testing;
  • drift testing;
  • side-effect testing;
  • reward-hacking testing where reinforcement/reward behavior is relevant;
  • API/interface testing;
  • ML functional-performance testing;
  • metamorphic testing;
  • back-to-back and A/B testing;
  • requirements and model-documentation review;
  • exploratory and fuzz testing for unexpected inputs;
  • performance testing;
  • smoke/regression tests around model updates;
  • red teaming for security, safety, privacy, or harmful-output risks.

The exam skill is to match the risk to the most appropriate test approach, not to select every technique.

Metamorphic K3 drill

Given: an OCR model should recognize a supported printed invoice independently of harmless image metadata.

Derive a source test case and at least three follow-up cases by:

  1. defining the source input;
  2. defining a valid transformation;
  3. stating the metamorphic relation between source and follow-up outputs;
  4. specifying any tolerance;
  5. explaining what a violation means.

Repeat for a second scenario where the expected relation is monotonic rather than invariant. This directly practices the K3 requirement to derive metamorphic test cases from a scenario.

Use these as visual reinforcement after reading the chapter. The ISTQB syllabus remains the exam authority.

Machine Learning Fundamentals: Bias and VarianceStatQuest with Josh Starmer · YouTube
SpeedEvery speed button sends the requested value to YouTube. Unsupported values such as 3× or 4× may be clamped by the embedded player.

Source registry

Reviewed 2026-09-06 · 4 chapter references
Certified Tester AI Testing Syllabus v2.0

Primary exam authority. Learning objectives, terminology, chapter scope, recommended training time, and hands-on objectives in this guide are mapped to this syllabus.

ISTQB · official syllabus
Source ↗
Model evaluation: quantifying the quality of predictions

Technical companion for classification metrics and model-evaluation examples. The exam definitions should still be learned from the ISTQB syllabus.

scikit-learn · technical reference
Source ↗
Artificial Intelligence Risk Management Framework

Practical companion for connecting AI quality, risk, governance, monitoring, and lifecycle controls to real systems.

NIST · risk framework
Source ↗
AI RMF: Generative Artificial Intelligence Profile

Additional risk vocabulary and mitigations for generative AI systems. Useful for GenAI/LLM test design and red-team exercises.

NIST · genai risk guidance
Source ↗