GimmeJob
Sign in
Certification learning path · CT-AI v2.0 · Chapter 09 / 12

ISTQB CT-AI v2.0 Exam Preparation

Hands-on labs & practical tasks

These labs are designed to reinforce the syllabus hands-on objectives. The certification exam is multiple-choice, but you should be able to execute the reasoning behind the questions rather than memorize vocabulary.

Lab 1 — Build the ML workflow

Goal: create, train and evaluate a small supervised model.

Use the Iris example from chapter 3 or another small public dataset. Produce:

  • problem statement and target;
  • train/test split rationale;
  • model configuration;
  • confusion matrix;
  • at least two suitable metrics;
  • one limitation;
  • one regression test you would keep for the next model version.

Done when: another tester can explain what evidence the test set provides and why it must remain separate from tuning.

Lab 2 — Prepare data and catch leakage

Create a deliberately dirty dataset with nulls, duplicates, invalid ranges, class imbalance and repeated entities. Write automated assertions for the constraints and design a split that prevents entity leakage.

Then introduce a leaked feature strongly correlated with the label. Observe how performance changes and explain why the “better” score is misleading.

Artifact: a data-quality checklist plus test code/query output.

Lab 3 — Classification-metric drill

For each matrix below calculate accuracy, precision, recall and F1 without a library.

A: TP=90, FP=10, FN=30, TN=870.

B: TP=18, FP=2, FN=2, TN=18.

C: TP=4, FP=1, FN=16, TN=979.

Then answer:

  • Which case shows why accuracy can be dangerous?
  • In which product would false negatives dominate risk?
  • How would changing a classification threshold affect precision/recall?

Verify calculations with a library only after doing them manually.

Lab 4 — Implement a perceptron calculation

Write a function that accepts features, weights and bias and returns a binary step output.

Python
editable · browser sandbox
Result
Run the code to see the result.

Create boundary cases where the weighted sum is just below, exactly at, and just above zero. Explain why boundary analysis remains useful even inside an AI-focused course.

Lab 5 — Exploratory LLM session

Test a model against one narrow policy requirement. Build a charter, 20+ prompt variations, a scoring rubric and a session log.

Required prompt groups:

  • normal/expected;
  • paraphrase;
  • missing evidence;
  • ambiguous intent;
  • prompt injection;
  • multi-turn escalation;
  • multilingual;
  • very long/noisy context.

Run at least five high-risk prompts three times each. Record whether variability changes the risk conclusion.

Artifact: test charter, transcript IDs, rubric, result summary, five regression prompts.

Lab 6 — Input-data test suite

Take a CSV or JSON dataset and implement at least ten checks covering:

  • schema;
  • types;
  • ranges;
  • allowed values;
  • nulls;
  • uniqueness;
  • duplicate entities;
  • class distribution;
  • subgroup coverage;
  • label sample review/provenance.

Add a deliberately broken pipeline transformation and prove a test catches it.

Lab 7 — Metamorphic model testing

Choose a classifier or ranking model where at least one relation should hold across transformed inputs. Define three metamorphic relations.

For each relation document:

relation:
precondition:
source input generation:
transformation:
expected output relation:
tolerance:
number of samples:
failure triage:

Execute the tests. A relation that turns out not to be valid is still useful if you can explain why the domain assumption was wrong and refine it.

Capstone — one complete AI test strategy

Choose one system: fraud classifier, medical triage, document classifier, recommendation engine, vision detector, or RAG assistant.

Your strategy must contain:

  1. intended use and boundaries;
  2. top ten product risks;
  3. AI quality characteristics and measurable acceptance criteria;
  4. input-data test plan;
  5. model test plan;
  6. system/integration tests;
  7. oracle strategy;
  8. adversarial/red-team tests;
  9. deployment evidence;
  10. production drift/performance monitoring;
  11. rollback/fallback;
  12. residual-risk owner.

If you can defend that strategy and explain why every test maps to a risk, you have converted the syllabus into working QA knowledge.

Source registry

Reviewed 2026-09-06 · 5 chapter references
Certified Tester AI Testing Syllabus v2.0

Primary exam authority. Learning objectives, terminology, chapter scope, recommended training time, and hands-on objectives in this guide are mapped to this syllabus.

ISTQB · official syllabus
Source ↗
Machine Learning Crash Course

Optional technical reinforcement for ML fundamentals, classification, neural networks, data, and generalization concepts used by CT-AI.

Google for Developers · technical learning
Source ↗
Model evaluation: quantifying the quality of predictions

Technical companion for classification metrics and model-evaluation examples. The exam definitions should still be learned from the ISTQB syllabus.

scikit-learn · technical reference
Source ↗
sklearn.metrics.confusion_matrix

Concrete implementation reference for confusion matrices used in the hands-on classification exercise.

scikit-learn · technical reference
Source ↗
AI RMF: Generative Artificial Intelligence Profile

Additional risk vocabulary and mitigations for generative AI systems. Useful for GenAI/LLM test design and red-team exercises.

NIST · genai risk guidance
Source ↗