GimmeJob
Sign in
Certification learning path · CT-AI v2.0 · Chapter 04 / 12

ISTQB CT-AI v2.0 Exam Preparation

3. Machine Learning

This is the largest syllabus chapter. The exam expects both conceptual understanding and application: you should be able to reason about the workflow, select data and evaluation approaches, and calculate classification metrics.

Forms of machine learning

Supervised learning learns from labeled examples. Typical tasks:

  • classification: spam/not spam, defect class, disease category;
  • regression: price, remaining useful life, demand.

Unsupervised learning works without target labels to discover structure, for example clustering or dimensionality reduction.

Reinforcement learning learns actions through interaction and rewards. Test design must consider policy behavior, exploration, environment assumptions, reward specification, and potentially long sequences of actions.

Semi/self-supervised approaches may use a mixture of limited labels and large unlabeled datasets. For the exam, focus on the distinctions and consequences described in the syllabus rather than collecting every modern ML taxonomy term.

The ML workflow

A practical lifecycle is:

  1. define the problem and measurable success criteria;
  2. acquire/select data;
  3. inspect and prepare data;
  4. split data appropriately;
  5. choose features/model/architecture;
  6. train;
  7. validate and tune;
  8. evaluate on held-out evidence;
  9. package and integrate;
  10. deploy;
  11. monitor data, model and system behavior;
  12. retrain or replace under controlled change.

Testing is not a final box. Test activities exist around data, pipeline code, model behavior, integrations, deployment, and monitoring.

Train, validation and test data

  • Training set: used to fit learned parameters.
  • Validation set: used during development for model selection, tuning and decisions.
  • Test set: held back for an unbiased final evaluation of the selected model.

A crucial failure is data leakage: information from the target, future, validation/test population, or duplicates leaks into training and creates unrealistic performance.

Example: randomly splitting rows from the same patient across train and test can let the model recognize patient-specific patterns. A group-aware split by patient may be required.

Pretrained models, fine-tuning and RAG

A pretrained/foundation model has already learned from a large source dataset. You can use it directly, adapt it through fine-tuning, or augment its input with retrieved information.

Fine-tuning changes model parameters using additional training data. Regression testing must cover the intended improvement and unintended capability/safety degradation.

Retrieval-Augmented Generation (RAG) retrieves external documents/chunks and supplies them as context to a generative model. Testing must separate:

  • retrieval quality: was relevant evidence found?
  • context construction: was the right content passed and safely delimited?
  • generation: did the model use the evidence correctly?
  • end-to-end answer quality: is the final response correct, grounded and appropriate?

Do not call RAG “training.” Retrieval changes runtime context; it does not by itself update model weights.

Data preparation

Common steps include validation, cleaning, deduplication, missing-value treatment, normalization/standardization, encoding categorical values, feature engineering, balancing/sampling, augmentation, and labeling.

Every transformation can introduce defects. Therefore the pipeline itself needs tests: schema, types, ranges, row counts, null rates, category mappings, deterministic transformations where expected, and train/serve consistency.

Confusion matrix

For a binary classifier:

  • TP: positive case predicted positive.
  • TN: negative case predicted negative.
  • FP: negative case incorrectly predicted positive.
  • FN: positive case incorrectly predicted negative.

From those values:

  • Accuracy = (TP + TN) / (TP + TN + FP + FN)
  • Precision = TP / (TP + FP)
  • Recall / sensitivity = TP / (TP + FN)
  • Specificity = TN / (TN + FP)
  • F1 = 2 × precision × recall / (precision + recall)

Worked example

Suppose TP=36, FP=9, FN=4, TN=51. Total = 100.

  • Accuracy = (36 + 51) / 100 = 0.87.
  • Precision = 36 / 45 = 0.80.
  • Recall = 36 / 40 = 0.90.
  • Specificity = 51 / 60 = 0.85.
  • F1 = 2 × 0.80 × 0.90 / 1.70 ≈ 0.847.

The correct metric depends on consequences. If a missed positive is dangerous, recall may dominate. If false alarms are extremely costly, precision or specificity may matter more. “Highest accuracy” is not an automatic answer.

Class imbalance trap

If only 1% of transactions are fraudulent, a classifier that predicts “not fraud” for every row is 99% accurate and useless. Always inspect class-specific metrics and business impact.

Thresholds change trade-offs

Many classifiers output a score/probability and use a threshold to produce a class. Moving the threshold changes FP/FN behavior. A threshold is therefore part of the product decision and test configuration, not just an internal model detail.

Test threshold behavior against explicit risk and acceptance criteria.

Neural networks

A simple artificial neuron combines inputs and weights, adds a bias, then applies an activation function. Training adjusts weights to reduce a loss/error signal.

inputs x weights -> weighted sum + bias -> activation -> output

Layers of neurons can learn progressively useful representations. Key testing implications include sensitivity to data, non-linear behavior, large input spaces, limited explainability, stochastic training, and the possibility that two training runs differ.

Perceptron intuition

For inputs x1 and x2 with weights w1 and w2 and bias b:

z = x1*w1 + x2*w2 + b
output = 1 if z >= 0 else 0

If x1=1, x2=0, w1=0.8, w2=-0.3, b=-0.2, then z=0.6 and the output is 1. You do not need deep calculus to solve this kind of reasoning task.

Neural-network coverage

Traditional code coverage does not tell you how broadly a neural network’s internal behavior has been exercised. Neural-network coverage criteria attempt to measure activation or structural behavior inside the network. They can help reveal unexercised behavior but do not prove correctness, safety, or complete input-space coverage.

Treat coverage as one signal for test adequacy, not a quality guarantee.

Hands-on: build and evaluate a tiny classifier

The point is to make the lifecycle concrete, not to memorize scikit-learn syntax.

from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix, classification_report
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=42, stratify=y
)

model = LogisticRegression(max_iter=500)
model.fit(X_train, y_train)
pred = model.predict(X_test)

print(confusion_matrix(y_test, pred))
print(classification_report(y_test, pred))

For the lab, identify which artifacts correspond to data, preprocessing, model, learned parameters, evaluation data and metric output. Then deliberately make one bad change—remove stratification, duplicate test rows into training, or reduce training data—and explain why the new score is or is not trustworthy.

Practice: without notes, calculate accuracy, precision, recall and F1 for TP=42, FP=14, FN=8, TN=136. Then write which metric you would prioritize for (a) cancer screening and (b) auto-blocking legitimate bank transfers, with one sentence explaining the cost of the relevant error.

Syllabus taxonomy you must be able to name

For AI-3.1.1, the core forms are supervised, unsupervised and reinforcement learning. Within them, remember the syllabus examples:

  • supervised → classification and ML regression;
  • unsupervised → clustering and association;
  • reinforcement learning → an agent learns through interaction, rewards and penalties.

Exact neural-network coverage measures

For AI-3.4.3, know what each named measure is trying to exercise:

  • Neuron coverage: proportion of neurons whose activation/output exceeds the chosen activation threshold during testing.
  • k-multisection neuron coverage (kMNC): divide a neuron's observed activation range into k sections and measure how many sections tests exercise.
  • Neuron boundary coverage (NBC): exercise neuron activations outside the lower/upper activation boundaries observed during training.

They are structural adequacy indicators, not proof of functional correctness or generalization.

Hands-on objective checklist for Chapter 3

The practical work in this guide maps to all Chapter 3 hands-on areas:

  • create an ML model;
  • perform data preparation supporting model creation;
  • evaluate a model with selected functional-performance metrics;
  • compare how different model/dataset combinations affect training and behavior;
  • experience a simple perceptron implementation.

Add one experiment to Lab 1: train at least two different model configurations or use two different train/test samples, compare their evaluation results, and explain why stochastic/data choices can change the observed behavior.

Use these as visual reinforcement after reading the chapter. The ISTQB syllabus remains the exam authority.

Machine Learning Fundamentals: The Confusion MatrixStatQuest with Josh Starmer · YouTube
SpeedEvery speed button sends the requested value to YouTube. Unsupported values such as 3× or 4× may be clamped by the embedded player.
But what is a neural network? | Deep learning chapter 13Blue1Brown · YouTube
SpeedEvery speed button sends the requested value to YouTube. Unsupported values such as 3× or 4× may be clamped by the embedded player.

Source registry

Reviewed 2026-09-06 · 4 chapter references
Certified Tester AI Testing Syllabus v2.0

Primary exam authority. Learning objectives, terminology, chapter scope, recommended training time, and hands-on objectives in this guide are mapped to this syllabus.

ISTQB · official syllabus
Source ↗
Machine Learning Crash Course

Optional technical reinforcement for ML fundamentals, classification, neural networks, data, and generalization concepts used by CT-AI.

Google for Developers · technical learning
Source ↗
Model evaluation: quantifying the quality of predictions

Technical companion for classification metrics and model-evaluation examples. The exam definitions should still be learned from the ISTQB syllabus.

scikit-learn · technical reference
Source ↗
sklearn.metrics.confusion_matrix

Concrete implementation reference for confusion matrices used in the hands-on classification exercise.

scikit-learn · technical reference
Source ↗