GimmeJob
Sign in
Certification learning path · CT-AI v2.0 · Chapter 06 / 12

ISTQB CT-AI v2.0 Exam Preparation

5. Input Data Testing

For ML systems, input data is part of the implementation. Bad data can produce a model that is perfectly trained to do the wrong thing.

Major input-data risks

Look for:

  • missing, malformed, corrupted or duplicated records;
  • wrong units, timestamps, encodings or categories;
  • target leakage or future information;
  • sampling bias and underrepresented operating conditions;
  • label noise or systematically wrong labels;
  • train/validation/test overlap;
  • data that is stale relative to production;
  • pipeline transformations that differ between training and serving;
  • personally identifiable or prohibited data included unexpectedly;
  • poisoned or adversarially inserted training data;
  • provenance/licensing/consent problems where relevant.

Mitigation is not only “clean the dataset.” It can include collection changes, stratification, re-labeling, constraints, pipeline tests, human review, weighting, augmentation, monitoring, access controls, and explicit limitation of the model’s operating domain.

Bias in data

Bias can enter through selection, measurement, historical processes, labels, proxies, missingness, or feedback loops.

A dataset can be numerically balanced and still biased. Example: equal counts by age group do not help if one group’s labels were generated by a systematically different process.

Test bias by asking:

  • Who/what is represented and who/what is missing?
  • Does measurement quality differ by group or condition?
  • Are protected or sensitive attributes used directly or through proxies?
  • Are labels based on past decisions that already contain bias?
  • Are subgroup sample sizes sufficient to support claims?
  • Does performance differ materially by slice?

Data pipeline testing

Treat ingestion and transformation as production code.

Test boundaries such as:

source -> ingestion -> schema validation -> cleaning -> feature transformation
       -> split -> training artifact
       -> serving transformation -> model input

Checks include:

  • schema and required fields;
  • type and range constraints;
  • allowed categorical values;
  • uniqueness and duplicate rate;
  • null/missing-value rate;
  • row counts before/after transformations;
  • distribution changes;
  • deterministic transformation where expected;
  • training/serving feature parity;
  • lineage/version metadata;
  • failure behavior when constraints are violated.

Classic defect: training divides a monetary field by 100 to convert cents to currency units, but serving sends cents directly. The model may appear defective when the real bug is pipeline inconsistency.

Representativeness

A dataset is representative when it adequately reflects the intended operational population and conditions for the claim being made.

Representativeness is contextual. A road-sign dataset collected only on sunny daytime roads is not representative for a system intended for night, rain and snow—even with millions of images.

Define operational dimensions such as geography, device, language, time, environment, subgroup, class, rarity and expected edge conditions. Then compare the dataset distribution against those dimensions.

Data constraints

Data constraints turn assumptions into executable checks.

Examples:

  • age between 0 and 120;
  • timestamp not in the future;
  • image width/height above minimum;
  • currency in supported set;
  • category belongs to the trained vocabulary;
  • no duplicate entity across train and test groups;
  • required field present for at least 99.9% of records;
  • class distribution within an agreed range.

A constraint violation does not always mean “delete the row.” It means the system should handle the condition deliberately.

Label correctness

Labels are ground truth only if the labeling process deserves trust.

Test label quality through:

  • random and risk-based sample review;
  • multiple annotators and agreement analysis;
  • expert adjudication for ambiguous cases;
  • clear labeling guidelines;
  • gold/reference examples;
  • targeted review of classes with low model performance;
  • checks for systematic disagreement by source or subgroup.

High disagreement can mean the labels are poor, but it can also expose an ambiguous requirement. That is product evidence, not just a data-cleaning problem.

Hands-on dataset audit

Create a small CSV with columns:

id, age, country, device, label
1, 34, UA, android, positive
2, 999, UA, android, positive
2, 28, PL, ios, negative
4, , DE, web, negative
5, 41, XX, android, maybe

Write checks for:

  1. unique id;
  2. age range;
  3. allowed countries/devices;
  4. missing values;
  5. label vocabulary;
  6. duplicate rows/entities;
  7. class balance by country and device.

Then extend the exercise: assume row id is a person and there are multiple observations per person. Explain why a row-level random train/test split may leak identity information and propose a group-level split.

Exam traps

  • Large dataset ≠ representative dataset.
  • Balanced classes ≠ unbiased labels.
  • Clean training data ≠ safe production inputs.
  • A pipeline schema test does not prove semantic correctness.
  • Labels should be tested, not blindly treated as truth.
  • Train/test leakage can happen through entities, time, duplicates or derived features even when files are separate.

Practice: for a speech-recognition model intended for all customers, design a data test matrix with at least six representativeness dimensions. For each dimension state how you would measure coverage and what you would do if the acceptance dataset has a gap.

Chapter 5 exam techniques to recognize

A complete answer to an input-data-risk scenario may involve more than schema assertions. CT-AI v2.0 includes approaches such as:

  • review of data sources, provenance and preparation;
  • exploratory data analysis (EDA);
  • static analysis/review of preparation or pipeline code;
  • dynamic testing of model outcomes across sensitive groups;
  • disparate impact analysis using realistic counterfactual changes to sensitive attributes and statistical comparison of outcomes;
  • data-pipeline testing at component, integration and system levels;
  • dataset constraint testing;
  • data representativeness testing;
  • label correctness testing;
  • multiple annotation and inter-annotator agreement as evidence about label reliability.

Disparate-impact drill

For a loan-decision model, create pairs of otherwise-valid applications where a relevant sensitive attribute changes while other decision-relevant facts remain controlled. Run enough cases to evaluate whether the outcome distribution changes materially. Before drawing a bias conclusion, verify the counterfactuals are realistic and that the sample is large enough to support the comparison.

This is stronger than changing one attribute in one record and declaring the system biased from a single output.

Source registry

Reviewed 2026-09-06 · 3 chapter references
Certified Tester AI Testing Syllabus v2.0

Primary exam authority. Learning objectives, terminology, chapter scope, recommended training time, and hands-on objectives in this guide are mapped to this syllabus.

ISTQB · official syllabus
Source ↗
Artificial Intelligence Risk Management Framework

Practical companion for connecting AI quality, risk, governance, monitoring, and lifecycle controls to real systems.

NIST · risk framework
Source ↗
Machine Learning Crash Course

Optional technical reinforcement for ML fundamentals, classification, neural networks, data, and generalization concepts used by CT-AI.

Google for Developers · technical learning
Source ↗