← All 186 courses
AI-212 AI & Data

Ray & Large-Scale Model Training

Engineer the training platform itself: Ray as the distributed substrate, DeepSpeed ZeRO and FSDP as the sharding layer, and Kubernetes as the scheduler underneath. Learners stand up a multi-node training job that survives node loss, tune it for utilisation rather than raw speed, and defend a run plan with cost, scaling-law and stop-criteria evidence rather than intuition.

B5Advanced
6modules
42skill atoms
1role journey
Curriculum

What this course covers.

6 modules, 42 named skill atoms. Expand any module to see them.

1Ray Core Concepts7 skill atoms
tasks and actorsobject store and zero-copy refsplacement groups for gang placementGPU resource requestsRay Dashboard inspectiondriver vs worker failure semanticsserialisation pitfalls on large objects
2Ray Train and Ray Tune7 skill atoms
TorchTrainer and ScalingConfigworker groups and per-worker setupcheckpoint reporting into the runASHA and population-based training schedulerssearch space definitiontrial concurrency vs GPU supplyearly stopping on a plateau
3DeepSpeed ZeRO and FSDP7 skill atoms
ZeRO stages 1 2 and 3optimizer state and parameter shardingCPU and NVMe offload costFSDP auto-wrap policiesfull vs sharded state dictmixed precision policy choiceZeRO vs FSDP selection criteria
4Multi-Node Scheduling on Kubernetes7 skill atoms
KubeRay operator and RayCluster CRDhead and worker node poolsgang scheduling and pending-pod deadlockGPU device plugin and node taintsshared storage for datasets and checkpointsNCCL and network env tuningspot interruption handling
5Fault Tolerance and Elastic Training7 skill atoms
worker failure retry policyresume from last durable checkpointelastic world-size changespreemption drain hookshung-rank detection and job killidempotent data sharding on restartblast radius of one bad node
6Cost, Utilisation and Scaling Laws7 skill atoms
GPU-hour cost per runMFU as the utilisation metriccompute-optimal sizing as a floor not a targetinference-cost-aware over-training past Chinchilla ratiostotal cost of ownership over the serving lifereserved vs spot capacity mixstop criteria and diminishing returns
Where it fits

AI-212 in a role journey.

This course appears in 1 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.

AI Infrastructure Engineer

Capstone stage
CL-405AI-212journey complete

Roles this course serves

The capability ladder

This course is authored to band B5.

Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.

What do B1–B6 mean?The CASI Capability Ladder — click to expand

Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.

What the learner can doTypical evidence
B1
AwareUnderstands concepts and vocabulary; uses tools with guidance
Knowledge checks
B2
FoundationPerforms standard tasks correctly in familiar contexts
Guided labs, autograded exercises
B3
PractitionerDelivers complete pieces of work independently
Scenario labs, proctored hands-on exams
B4
ProfessionalHandles production-grade complexity, trade-offs and failure modes
Break-fix drills, design defenses
B5
AdvancedEngineers systems end-to-end under constraints; leads others
Rubric-scored capstones, vivas
B6
ExpertSets direction; recognised authority across teams
Portfolio + panel evaluation

A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.

Next step

Run AI-212 for your team.

This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.

Add it to a training plan Talk to our team Check your team’s level free