← All 186 courses
AI-214 AI & Data

Synthetic Data Generation & Evaluation

Use models to manufacture training and evaluation data without poisoning the thing you are training. Learners design seeds that produce genuine diversity, distil from a stronger teacher under real licence constraints, and build the filtering, dedup and drift checks that separate a usable synthetic corpus from an expensive echo of the generator.

B4Professional
5modules
35skill atoms
3role journeys
Curriculum

What this course covers.

5 modules, 35 named skill atoms. Expand any module to see them.

1When Synthetic Data Earns Its Place7 skill atoms
data scarcity vs class imbalance casescold-start eval set creationaugmenting rare edge casescost per generated samplehuman-labelled validation floortasks where synthesis fails outrightsuccess criteria fixed before generation
2Seed and Prompt Design for Diversity7 skill atoms
seed taxonomy and topic gridspersona and scenario conditioningSelf-Instruct and Evol-Instruct patternstemperature and top_p for spreadmode collapse from a single templatedifficulty ladderingschema-constrained generation
3Distillation from Stronger Models7 skill atoms
teacher output and rationale capturerejection sampling on verified answersmulti-teacher blendinglicence terms on model outputsteacher error propagation into the studentstudent capability ceilingcost per accepted sample
4Quality Filtering and Dedup7 skill atoms
heuristic and length filtersMinHash and LSH near-duplicate removalembedding-cluster dedupLLM-judge quality scoringtoxicity and PII screeningcontamination check against eval setsaccept-rate tracking per batch
5Privacy, Collapse and Provenance7 skill atoms
differential privacy budgets in synthesismembership inference riskre-identification via quasi-identifiersmodel collapse from recursive trainingdistribution drift against the real corpusKL and embedding-distance drift checksprovenance manifest and dataset card
Where it fits

AI-214 in the role journeys.

This course appears in 3 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.

AI / ML Engineer

Professional stage
AI-207AI-214AI-217

Data Scientist

Professional stage
AI-210AI-214AI-217

AI Quality / Eval Engineer

Professional stage
AI-404AI-214AI-509

Roles this course serves

The capability ladder

This course is authored to band B4.

Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.

What do B1–B6 mean?The CASI Capability Ladder — click to expand

Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.

What the learner can doTypical evidence
B1
AwareUnderstands concepts and vocabulary; uses tools with guidance
Knowledge checks
B2
FoundationPerforms standard tasks correctly in familiar contexts
Guided labs, autograded exercises
B3
PractitionerDelivers complete pieces of work independently
Scenario labs, proctored hands-on exams
B4
ProfessionalHandles production-grade complexity, trade-offs and failure modes
Break-fix drills, design defenses
B5
AdvancedEngineers systems end-to-end under constraints; leads others
Rubric-scored capstones, vivas
B6
ExpertSets direction; recognised authority across teams
Portfolio + panel evaluation

A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.

Next step

Run AI-214 for your team.

This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.

Add it to a training plan Talk to our team Check your team’s level free