Final review, references & exam-day checklist
Use this section after the learning chapters, not as a substitute for them.
High-yield distinctions
Be able to explain each pair in one or two sentences:
- conventional deterministic logic vs learned/probabilistic behavior;
- locked vs adaptive model;
- training vs validation vs test dataset;
- data cleaning vs proving representativeness;
- global performance vs subgroup/slice performance;
- precision vs recall;
- test oracle vs one reference example;
- exact assertion vs statistical/rubric/metamorphic oracle;
- input-data testing vs model testing vs system testing;
- adversarial testing vs ordinary negative testing;
- metamorphic vs back-to-back testing;
- A/B vs back-to-back testing;
- data drift vs concept drift;
- overfitting vs underfitting;
- fine-tuning vs RAG;
- model quality vs complete system safety.
Formula sheet
For binary classification:
accuracy = (TP + TN) / (TP + TN + FP + FN)
precision = TP / (TP + FP)
recall = TP / (TP + FN)
specificity = TN / (TN + FP)
F1 = 2 * precision * recall / (precision + recall)Before calculating, write what the positive class means. A formula can be correct while the interpretation is reversed.
Technique trigger sheet
When the scenario says… think first about:
- “No exact expected output” → oracle alternatives, metamorphic, statistical/rubric evaluation.
- “Small safe transformation should preserve behavior” → metamorphic testing.
- “Compare old and new model on same inputs” → back-to-back testing.
- “Different live groups receive alternatives” → A/B testing.
- “Input distribution changed” → drift and refreshed evidence.
- “Training great, validation worse” → overfitting or split/leakage/data issue.
- “Both training and validation poor” → underfitting/data/problem formulation.
- “Rare positive, missing it is costly” → recall/false negatives.
- “False alarms are costly” → precision/specificity/false positives depending on the requirement.
- “LLM safety bypass” → red teaming/adversarial exploration plus regression corpus.
- “Wrong model in production” → ML development/deployment artifact traceability.
- “Overall score looks good, subgroup bad” → slicing/fairness/representativeness and risk-specific acceptance criteria.
Official-material sequence
- Read the current CT-AI v2.0 certification page for exam logistics and version status.
- Keep the v2.0 syllabus open while studying; every weakness should map back to a learning objective.
- Use the ISTQB glossary for terminology disputes.
- Take the official Sample Exam Questions v2.2 under timed conditions.
- Use the official Sample Exam Answers v2.2 to review reasoning.
- Re-study weak objectives, then take this guide’s original mock.
Never rely on an old CT-AI v1.0 course without checking the version. v2.0 reorganized the syllabus into seven chapters, adds dedicated GenAI/LLM testing and removes the “using AI for testing” scope.
24-hour checklist
- I can calculate accuracy, precision, recall and F1 by hand.
- I can explain why accuracy can fail on imbalanced data.
- I can identify leakage scenarios.
- I can turn a vague AI quality claim into a measurable acceptance criterion.
- I can distinguish locked and adaptive systems.
- I can name oracle strategies for non-deterministic output.
- I can design a red-team charter for an LLM.
- I can propose representative data slices and label-quality checks.
- I can create at least two valid metamorphic relations.
- I can distinguish A/B and back-to-back testing.
- I can explain drift, overfitting and underfitting.
- I can describe deployment evidence tying the approved model to the serving artifact.
If any item is “no,” revise that topic instead of re-reading everything.
Exam execution
- Read the qualifier words: best, most appropriate, first, directly, likely.
- Identify the lifecycle stage and test object before choosing a technique.
- For numeric questions, write TP/TN/FP/FN explicitly before applying a formula.
- Eliminate answers that claim one metric or technique “proves” overall quality.
- Flag uncertain items and return; do not spend a large fraction of the exam on one question.
- Reserve final minutes to review flagged answers and accidental misreads.
After certification
CT-AI is a knowledge baseline, not the end state. Keep the capstone strategy from this guide and apply it to a real AI feature. Add production monitoring, evaluation datasets, red-team regressions, and deployment traceability. That work turns a certificate into demonstrable AI-testing capability.
Final practice: explain the complete lifecycle aloud in five minutes: requirements/quality → data → model training/evaluation → system testing → deployment → monitoring/change. If you cannot connect a technique to the risk it reduces, revisit that section before the exam.