← All 186 courses
AI-213 AI & Data

Preference Tuning: RLHF, DPO & Reward Models

Shape model behaviour after pre-training with the alignment stack teams actually run: preference data you can trust, reward models that generalise, and the choice between PPO-style RLHF and the simpler DPO family. Learners train a reward model, run both a PPO and a DPO variant on the same preference set, and prove the aligned model gained behaviour without quietly losing capability.

B5Advanced
6modules
42skill atoms
3role journeys
Curriculum

What this course covers.

6 modules, 42 named skill atoms. Expand any module to see them.

1Alignment After Pre-Training7 skill atoms
SFT vs preference tuning roleshelpfulness and harmlessness objectivesinstruction data quality floorbase capability as the alignment ceilingpolicy and reference model rolesKL constraint intuitionpipeline order and dependencies
2Preference Data Collection7 skill atoms
pairwise comparison interfacesannotation rubric designinter-annotator agreement with Cohen kappatie and both-bad handlingprompt distribution coverageAI feedback vs human labelstolerable label-noise budget
3Reward Model Training7 skill atoms
Bradley-Terry pairwise lossinitialising from the SFT checkpointheld-out preference accuracylength bias in reward scoresreward margin calibrationensembles for uncertaintyout-of-distribution reward blowup
4Policy-Gradient RL: PPO and GRPO7 skill atoms
rollout generation and scoringadvantage estimation with GAEGRPO group-relative advantage without a value headverifiable rewards (RLVR) on checkable tasksKL penalty against the reference policyTRL PPOTrainer and GRPOTrainer mechanicscompute cost relative to SFT
5DPO and Its Variants7 skill atoms
implicit reward derivationbeta temperature choiceIPO and KTO and ORPO differenceson-policy vs offline preference datareference-free variantsTRL DPOTrainer setupwhere DPO underperforms PPO
6Over-Optimisation and Evaluation7 skill atoms
reward hacking signaturesverbosity and sycophancy driftGoodhart failure on a proxy rewardcapability regression on held-out benchmarkswin-rate judging with position-bias controlKL vs reward frontier plotsrelease gate and rollback criteria
Where it fits

AI-213 in the role journeys.

This course appears in 3 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.

AI / ML Engineer

Capstone stage
AI-510AI-213journey complete

GenAI / LLM Engineer

Capstone stage
AI-510AI-213journey complete

AI Quality / Eval Engineer

Capstone stage
AI-402AI-213journey complete

Roles this course serves

The capability ladder

This course is authored to band B5.

Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.

What do B1–B6 mean?The CASI Capability Ladder — click to expand

Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.

What the learner can doTypical evidence
B1
AwareUnderstands concepts and vocabulary; uses tools with guidance
Knowledge checks
B2
FoundationPerforms standard tasks correctly in familiar contexts
Guided labs, autograded exercises
B3
PractitionerDelivers complete pieces of work independently
Scenario labs, proctored hands-on exams
B4
ProfessionalHandles production-grade complexity, trade-offs and failure modes
Break-fix drills, design defenses
B5
AdvancedEngineers systems end-to-end under constraints; leads others
Rubric-scored capstones, vivas
B6
ExpertSets direction; recognised authority across teams
Portfolio + panel evaluation

A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.

Next step

Run AI-213 for your team.

This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.

Add it to a training plan Talk to our team Check your team’s level free