1. Introduction to Artificial Intelligence
The exam does not require you to become a data scientist. It does require a tester to understand what makes an AI-based system materially different from conventional deterministic software and how those differences change test strategy.
AI-based systems versus conventional systems
Conventional software usually implements behavior explicitly through code and rules. Given the same state and input, deterministic code should normally produce the same output. AI-based behavior may instead be learned from data and may be probabilistic. The implementation therefore includes more than source code: training data, validation data, model architecture, learned parameters, preprocessing, configuration, inference runtime, and often external foundation models.
Testing consequence: a passing unit test for application code does not establish that the learned behavior is correct, representative, robust, safe, or stable after data changes.
A useful decomposition is:
- Conventional part: routing, API contracts, permissions, persistence, UI, deterministic calculations.
- AI part: feature extraction, model inference, ranking/classification/generation, confidence or probability outputs.
- Integration boundary: how deterministic application logic interprets uncertain model output.
Example: an expense app uses a model to classify receipt images. “File upload accepts JPEG” is conventional behavior. “A restaurant receipt is categorized as meals with acceptable error rates across languages and lighting conditions” is AI behavior.
Narrow, general and super AI
For exam purposes, distinguish scope rather than marketing claims:
- Narrow AI: designed for a bounded task or domain, such as fraud scoring, object detection, translation, or recommendation.
- Artificial General Intelligence (AGI): hypothetical/general capability across a broad range of cognitive tasks comparable to general human intelligence.
- Super AI: hypothetical capability exceeding human intelligence broadly.
Most real systems you test today are narrow AI, even when a foundation model can perform many tasks.
Families of AI technology
Recognize common technology families and the testing implications they introduce:
- Machine learning: behavior learned from data rather than encoded entirely as rules.
- Deep learning / neural networks: multi-layer learned representations; can be powerful but difficult to explain and highly data-dependent.
- Natural-language processing: text understanding or generation; introduces ambiguity, language variation, semantic equivalence, and harmful-content risks.
- Computer vision: image/video inputs; sensitive to lighting, viewpoint, occlusion, resolution, and adversarial perturbation.
- Robotics/autonomous systems: perception plus action; safety, real-time constraints, environment interaction, sensors, and fallback behavior become central.
- Knowledge/rule-based approaches: more explicit logic; often easier to inspect but still can combine with learned components.
Do not equate “AI” with “LLM.” CT-AI covers both traditional ML and generative AI.
Generative AI
Generative AI produces new content such as text, images, audio, video, or code. Large language models generate sequences probabilistically from learned patterns. Important testing consequences include:
- different valid outputs for the same prompt;
- hallucination or unsupported claims;
- sensitivity to prompt wording and context;
- safety-policy bypasses;
- prompt injection and untrusted retrieved content;
- quality that is multidimensional rather than a single right/wrong result.
A tester therefore needs properties, rubrics, statistical samples, reference sets, human or model-assisted evaluation, and red-team scenarios rather than only exact-string assertions.
Hardware matters
Model training and inference can depend on CPUs, GPUs, TPUs/AI accelerators, memory, storage bandwidth, and device-specific runtimes. The hardware choice can affect latency, throughput, energy use, numerical precision, supported operations, and sometimes model outputs.
Testing example: a vision model validated on a datacenter GPU is converted to a lower-precision mobile format. Re-test functional performance, latency, memory, thermal behavior, and representative devices after conversion. Treat the converted artifact as a new test object, not merely a deployment detail.
Developing and hosting AI models
A model can be:
- built and trained internally;
- fine-tuned from a pretrained model;
- consumed as a third-party hosted API;
- deployed to your own cloud/runtime;
- executed on-device or at the edge.
Each option changes controllability and observability. With a third-party API you may not control model weights or update timing, so contract testing, version pinning where available, monitoring, fallback behavior, and regression datasets become more important.
ML development frameworks
Frameworks such as TensorFlow, PyTorch, scikit-learn and related tooling provide model construction, training, evaluation and serialization. A tester does not need to memorize APIs, but should understand that framework/runtime/library versions are part of reproducibility and deployment risk.
Two model files with identical names are not enough evidence of identical behavior if preprocessing code, dependency versions, random seeds, hardware kernels, or data snapshots differ.
Regulations, standards and governance
AI systems can be subject to legislation, sector rules, contractual controls and technical standards. The tester’s role is not to give legal advice; it is to translate applicable obligations and quality expectations into testable acceptance criteria, evidence, traceability, monitoring, and release controls.
For this syllabus, ISO/IEC 25059 is especially important because it provides an AI-system quality model. NIST AI RMF is a useful practical companion for risk thinking. Regulations may evolve faster than certification syllabi, so separate “what the exam expects” from “what the current law requires in a specific deployment.”
Exam traps
- Assuming every AI system is adaptive in production. Many deployed models are locked until deliberately replaced.
- Calling an LLM “general AI” simply because it supports many tasks.
- Testing only application code and ignoring data/model artifacts.
- Treating a hosted model API as risk-free because another company owns the model.
- Confusing a probabilistic result with a random or untestable result.
Practice: choose one AI feature you know. Draw its boundary with four boxes: input/data processing, model, application logic, external dependencies. Under each box list two risks and one test. Then state whether the model is locked or adaptive and what evidence would prove your answer.
CT-AI terminology checkpoint
For the exam, keep the syllabus distinctions precise:
- Narrow AI is the category for deployed task/domain-focused AI systems.
- Frontier AI is discussed as advanced state-of-the-art AI while still not being equivalent to general AI.
- General AI and super AI are distinct concepts from the narrow AI systems deployed today.
- When the syllabus says ML regression, it means prediction of continuous values. Do not confuse this with software regression testing after a change.
Use this vocabulary when a question asks you to classify an AI system rather than substituting looser product-marketing terms.