Preference Tuning: RLHF, DPO & Reward Models
Shape model behaviour after pre-training with the alignment stack teams actually run: preference data you can trust, reward models that generalise, and the choice between PPO-style RLHF and the simpler DPO family. Learners train a reward model, run both a PPO and a DPO variant on the same preference set, and prove the aligned model gained behaviour without quietly losing capability.
What this course covers.
6 modules, 42 named skill atoms. Expand any module to see them.
1Alignment After Pre-Training7 skill atoms
2Preference Data Collection7 skill atoms
3Reward Model Training7 skill atoms
4Policy-Gradient RL: PPO and GRPO7 skill atoms
5DPO and Its Variants7 skill atoms
6Over-Optimisation and Evaluation7 skill atoms
AI-213 in the role journeys.
This course appears in 3 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.
Roles this course serves
This course is authored to band B5.
Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.
What do B1–B6 mean?The CASI Capability Ladder — click to expand
Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.
A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.
Other AI courses at this level.
Model Compression: Quantization, PEFT, QLoRA & Distillation
GPU Engineering for AI (CUDA · cuDNN · TensorRT)
Ray & Large-Scale Model Training
Knowledge Graphs & GraphRAG
Multi-Agent Systems & Orchestration
LLM Inference Optimization & Serving (vLLM · LiteLLM)
Run AI-213 for your team.
This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.