Hands-on labs & practical tasks
These labs are designed to reinforce the syllabus hands-on objectives. The certification exam is multiple-choice, but you should be able to execute the reasoning behind the questions rather than memorize vocabulary.
Lab 1 — Build the ML workflow
Goal: create, train and evaluate a small supervised model.
Use the Iris example from chapter 3 or another small public dataset. Produce:
- problem statement and target;
- train/test split rationale;
- model configuration;
- confusion matrix;
- at least two suitable metrics;
- one limitation;
- one regression test you would keep for the next model version.
Done when: another tester can explain what evidence the test set provides and why it must remain separate from tuning.
Lab 2 — Prepare data and catch leakage
Create a deliberately dirty dataset with nulls, duplicates, invalid ranges, class imbalance and repeated entities. Write automated assertions for the constraints and design a split that prevents entity leakage.
Then introduce a leaked feature strongly correlated with the label. Observe how performance changes and explain why the “better” score is misleading.
Artifact: a data-quality checklist plus test code/query output.
Lab 3 — Classification-metric drill
For each matrix below calculate accuracy, precision, recall and F1 without a library.
A: TP=90, FP=10, FN=30, TN=870.
B: TP=18, FP=2, FN=2, TN=18.
C: TP=4, FP=1, FN=16, TN=979.
Then answer:
- Which case shows why accuracy can be dangerous?
- In which product would false negatives dominate risk?
- How would changing a classification threshold affect precision/recall?
Verify calculations with a library only after doing them manually.
Lab 4 — Implement a perceptron calculation
Write a function that accepts features, weights and bias and returns a binary step output.
Run the code to see the result.
Create boundary cases where the weighted sum is just below, exactly at, and just above zero. Explain why boundary analysis remains useful even inside an AI-focused course.
Lab 5 — Exploratory LLM session
Test a model against one narrow policy requirement. Build a charter, 20+ prompt variations, a scoring rubric and a session log.
Required prompt groups:
- normal/expected;
- paraphrase;
- missing evidence;
- ambiguous intent;
- prompt injection;
- multi-turn escalation;
- multilingual;
- very long/noisy context.
Run at least five high-risk prompts three times each. Record whether variability changes the risk conclusion.
Artifact: test charter, transcript IDs, rubric, result summary, five regression prompts.
Lab 6 — Input-data test suite
Take a CSV or JSON dataset and implement at least ten checks covering:
- schema;
- types;
- ranges;
- allowed values;
- nulls;
- uniqueness;
- duplicate entities;
- class distribution;
- subgroup coverage;
- label sample review/provenance.
Add a deliberately broken pipeline transformation and prove a test catches it.
Lab 7 — Metamorphic model testing
Choose a classifier or ranking model where at least one relation should hold across transformed inputs. Define three metamorphic relations.
For each relation document:
relation:
precondition:
source input generation:
transformation:
expected output relation:
tolerance:
number of samples:
failure triage:Execute the tests. A relation that turns out not to be valid is still useful if you can explain why the domain assumption was wrong and refine it.
Capstone — one complete AI test strategy
Choose one system: fraud classifier, medical triage, document classifier, recommendation engine, vision detector, or RAG assistant.
Your strategy must contain:
- intended use and boundaries;
- top ten product risks;
- AI quality characteristics and measurable acceptance criteria;
- input-data test plan;
- model test plan;
- system/integration tests;
- oracle strategy;
- adversarial/red-team tests;
- deployment evidence;
- production drift/performance monitoring;
- rollback/fallback;
- residual-risk owner.
If you can defend that strategy and explain why every test maps to a risk, you have converted the syllabus into working QA knowledge.