GimmeJob
Sign in
Certification learning path · CT-AI v2.0 · Chapter 03 / 12

ISTQB CT-AI v2.0 Exam Preparation

2. AI quality & acceptance criteria

A model can be highly accurate and still be an unacceptable product. CT-AI expects you to reason about quality characteristics that become especially important for AI systems and to turn vague statements such as “the model must be fair” into measurable test conditions.

AI-specific quality characteristics

Learn the characteristics in the wording and grouping used by the current syllabus/ISO model. More importantly, understand how each becomes observable evidence. Typical AI-specific concerns include behavior around:

  • functional adaptability / ability to cope with changing conditions where relevant;
  • transparency and explainability of behavior and decisions;
  • controllability and the ability for authorized humans/systems to direct, constrain, stop, or override behavior;
  • robustness against variations, noise, unusual inputs, and attacks;
  • fairness / freedom from inappropriate bias across relevant groups or conditions;
  • data quality and representativeness because learned behavior depends on data;
  • safety where model behavior can contribute to harm.

Use the official syllabus as the exam vocabulary source. Standards evolve, so do not substitute a newer terminology list for the version the exam is based on.

Quality characteristics interact

Quality goals can conflict. Maximizing one measure in isolation may reduce another.

Example: a fraud model can lower the decision threshold to catch more fraud. Recall rises, but false positives may block legitimate customers. The correct acceptance criteria depend on business loss, customer harm, review capacity, and safety/compliance constraints.

Example: an LLM may be constrained to refuse uncertain medical advice. That can reduce harmful responses but can also reduce usefulness. Test both safety and task effectiveness; neither metric alone proves acceptable quality.

AI and safety

Safety analysis starts with consequences, not model architecture.

Ask:

  1. What hazardous outcome can the AI contribute to?
  2. Who or what can be harmed?
  3. Under which operational conditions does risk increase?
  4. What preventive, detective and recovery controls exist?
  5. Which controls are outside the model itself?

For safety-related systems, test the complete control loop: model output, application interpretation, human override, safe-state behavior, alarms, logging, and recovery. A high model score cannot compensate for an unsafe integration decision.

Write measurable acceptance criteria

Bad criterion: “The model should be accurate and unbiased.”

Better criterion:

  • On the frozen acceptance dataset, emergency-case recall shall be at least 0.97.
  • Recall shall be measured separately for each clinically relevant subgroup with at least the agreed sample size.
  • No subgroup recall shall be more than 0.03 below the overall recall without documented risk acceptance.
  • The 95th-percentile inference latency on the target device shall stay below 250 ms.
  • When confidence is below the operational threshold, the system shall route the case to manual review rather than auto-decide.

Now every statement points to a test object, dataset, metric, threshold, environment, and expected behavior.

Acceptance-criteria checklist

For each criterion identify:

  • population / operating domain being claimed;
  • dataset or live sample used for measurement;
  • metric and exact calculation;
  • threshold and confidence/tolerance where relevant;
  • subgroups / slices that need separate evidence;
  • hardware/runtime/model version under test;
  • fallback behavior outside the acceptable region;
  • owner who accepts residual risk.

Exam traps

  • “Accuracy ≥ 95%” without specifying the dataset or population is not enough.
  • A global metric can hide failures in rare classes or subgroups.
  • Explainability is not the same as correctness.
  • Fairness is not automatically proven by equal overall accuracy.
  • Safety is a system property; do not restrict safety testing to the model component.

Practice: rewrite these three requirements as measurable criteria: “recommendations must be fair,” “the chatbot must be safe,” and “the detector must be reliable.” For each, include dataset/population, metric, threshold, slice, fallback, and environment.

Exact Chapter 2 exam vocabulary

For AI-2.1.1, learn the quality-characteristic names used by CT-AI v2.0. The syllabus discusses these new or adapted ISO/IEC 25059 characteristics:

  • AI functional correctness — acceptable correctness must be expressed with measurable error/performance criteria for probabilistic behavior.
  • Functional adaptability — the system's ability to adapt to changes in its operational environment where adaptation is part of the design.
  • User controllability — a human or external agent can influence/control the AI system appropriately.
  • Transparency — stakeholders receive appropriate information about the system, its data, behavior, limitations, or decisions.
  • AI robustness — acceptable behavior is maintained under relevant variations, disturbances, defects, or hostile conditions.
  • Intervenability — authorized intervention can stop, alter, or constrain operation when required.
  • Societal and ethical risk mitigation — the system mitigates unacceptable societal/ethical risks such as discriminatory or harmful behavior.

Safety is also a Chapter 2 keyword and has its own learning objective about the special considerations of safety-related AI systems.

Fairness, data quality, explainability, privacy and similar concerns are important test topics, but in an exam question asking specifically for the Chapter 2 / ISO/IEC 25059 quality-characteristic classification, do not replace the syllabus names above with a generic AI-trustworthiness list.

Classification drill

Classify the primary quality characteristic in each case before looking at any notes:

  1. A human operator must be able to cancel an AI-proposed critical action before execution.
  2. An adaptive control system must continue meeting agreed behavior after a supported environmental change.
  3. A classifier may make errors, but class-specific error rates must remain within approved limits.
  4. Documentation must expose the model version, intended use, known limitations and data provenance.
  5. Small supported input perturbations must not cause unacceptable behavior changes.

Then explain why one scenario can involve more than one characteristic even when the question asks for the best classification.

Source registry

Reviewed 2026-09-06 · 3 chapter references
Certified Tester AI Testing Syllabus v2.0

Primary exam authority. Learning objectives, terminology, chapter scope, recommended training time, and hands-on objectives in this guide are mapped to this syllabus.

ISTQB · official syllabus
Source ↗
ISO/IEC 25059:2023 — Quality model for AI systems

AI-system quality model referenced by the CT-AI syllabus. Use it to understand the source of AI-specific quality characteristics; exam wording follows the syllabus.

ISO · standard
Source ↗
Artificial Intelligence Risk Management Framework

Practical companion for connecting AI quality, risk, governance, monitoring, and lifecycle controls to real systems.

NIST · risk framework
Source ↗