5. Input Data Testing
For ML systems, input data is part of the implementation. Bad data can produce a model that is perfectly trained to do the wrong thing.
Major input-data risks
Look for:
- missing, malformed, corrupted or duplicated records;
- wrong units, timestamps, encodings or categories;
- target leakage or future information;
- sampling bias and underrepresented operating conditions;
- label noise or systematically wrong labels;
- train/validation/test overlap;
- data that is stale relative to production;
- pipeline transformations that differ between training and serving;
- personally identifiable or prohibited data included unexpectedly;
- poisoned or adversarially inserted training data;
- provenance/licensing/consent problems where relevant.
Mitigation is not only “clean the dataset.” It can include collection changes, stratification, re-labeling, constraints, pipeline tests, human review, weighting, augmentation, monitoring, access controls, and explicit limitation of the model’s operating domain.
Bias in data
Bias can enter through selection, measurement, historical processes, labels, proxies, missingness, or feedback loops.
A dataset can be numerically balanced and still biased. Example: equal counts by age group do not help if one group’s labels were generated by a systematically different process.
Test bias by asking:
- Who/what is represented and who/what is missing?
- Does measurement quality differ by group or condition?
- Are protected or sensitive attributes used directly or through proxies?
- Are labels based on past decisions that already contain bias?
- Are subgroup sample sizes sufficient to support claims?
- Does performance differ materially by slice?
Data pipeline testing
Treat ingestion and transformation as production code.
Test boundaries such as:
source -> ingestion -> schema validation -> cleaning -> feature transformation
-> split -> training artifact
-> serving transformation -> model inputChecks include:
- schema and required fields;
- type and range constraints;
- allowed categorical values;
- uniqueness and duplicate rate;
- null/missing-value rate;
- row counts before/after transformations;
- distribution changes;
- deterministic transformation where expected;
- training/serving feature parity;
- lineage/version metadata;
- failure behavior when constraints are violated.
Classic defect: training divides a monetary field by 100 to convert cents to currency units, but serving sends cents directly. The model may appear defective when the real bug is pipeline inconsistency.
Representativeness
A dataset is representative when it adequately reflects the intended operational population and conditions for the claim being made.
Representativeness is contextual. A road-sign dataset collected only on sunny daytime roads is not representative for a system intended for night, rain and snow—even with millions of images.
Define operational dimensions such as geography, device, language, time, environment, subgroup, class, rarity and expected edge conditions. Then compare the dataset distribution against those dimensions.
Data constraints
Data constraints turn assumptions into executable checks.
Examples:
- age between 0 and 120;
- timestamp not in the future;
- image width/height above minimum;
- currency in supported set;
- category belongs to the trained vocabulary;
- no duplicate entity across train and test groups;
- required field present for at least 99.9% of records;
- class distribution within an agreed range.
A constraint violation does not always mean “delete the row.” It means the system should handle the condition deliberately.
Label correctness
Labels are ground truth only if the labeling process deserves trust.
Test label quality through:
- random and risk-based sample review;
- multiple annotators and agreement analysis;
- expert adjudication for ambiguous cases;
- clear labeling guidelines;
- gold/reference examples;
- targeted review of classes with low model performance;
- checks for systematic disagreement by source or subgroup.
High disagreement can mean the labels are poor, but it can also expose an ambiguous requirement. That is product evidence, not just a data-cleaning problem.
Hands-on dataset audit
Create a small CSV with columns:
id, age, country, device, label
1, 34, UA, android, positive
2, 999, UA, android, positive
2, 28, PL, ios, negative
4, , DE, web, negative
5, 41, XX, android, maybeWrite checks for:
- unique id;
- age range;
- allowed countries/devices;
- missing values;
- label vocabulary;
- duplicate rows/entities;
- class balance by country and device.
Then extend the exercise: assume row id is a person and there are multiple observations per person. Explain why a row-level random train/test split may leak identity information and propose a group-level split.
Exam traps
- Large dataset ≠ representative dataset.
- Balanced classes ≠ unbiased labels.
- Clean training data ≠ safe production inputs.
- A pipeline schema test does not prove semantic correctness.
- Labels should be tested, not blindly treated as truth.
- Train/test leakage can happen through entities, time, duplicates or derived features even when files are separate.
Practice: for a speech-recognition model intended for all customers, design a data test matrix with at least six representativeness dimensions. For each dimension state how you would measure coverage and what you would do if the acceptance dataset has a gap.
Chapter 5 exam techniques to recognize
A complete answer to an input-data-risk scenario may involve more than schema assertions. CT-AI v2.0 includes approaches such as:
- review of data sources, provenance and preparation;
- exploratory data analysis (EDA);
- static analysis/review of preparation or pipeline code;
- dynamic testing of model outcomes across sensitive groups;
- disparate impact analysis using realistic counterfactual changes to sensitive attributes and statistical comparison of outcomes;
- data-pipeline testing at component, integration and system levels;
- dataset constraint testing;
- data representativeness testing;
- label correctness testing;
- multiple annotation and inter-annotator agreement as evidence about label reliability.
Disparate-impact drill
For a loan-decision model, create pairs of otherwise-valid applications where a relevant sensitive attribute changes while other decision-relevant facts remain controlled. Run enough cases to evaluate whether the outcome distribution changes materially. Before drawing a bias conclusion, verify the counterfactuals are realistic and that the sample is large enough to support the comparison.
This is stronger than changing one attribute in one record and declaring the system biased from a single output.