Model Compression: Quantization, PEFT, QLoRA & Distillation
Make big models economically deployable: quantization formats and their quality cost, parameter-efficient fine-tuning with LoRA/QLoRA, and distillation of large-model behaviour into small servable models — anchored in a cost-per-request business case, not just benchmarks.
What this course covers.
5 modules, 31 named skill atoms. Expand any module to see them.
1Why Compression Pays7 skill atoms
2Quantization in Practice7 skill atoms
3LoRA & QLoRA6 skill atoms
4Distillation6 skill atoms
5₹-per-Request Verdict5 skill atoms
AI-208 in a role journey.
This course appears in 1 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.
Roles this course serves
This course is authored to band B5.
Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.
What do B1–B6 mean?The CASI Capability Ladder — click to expand
Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.
A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.
Other AI courses at this level.
GPU Engineering for AI (CUDA · cuDNN · TensorRT)
Ray & Large-Scale Model Training
Preference Tuning: RLHF, DPO & Reward Models
Knowledge Graphs & GraphRAG
Multi-Agent Systems & Orchestration
LLM Inference Optimization & Serving (vLLM · LiteLLM)
Run AI-208 for your team.
This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.