GPU Engineering for AI (CUDA · cuDNN · TensorRT)
Understand and exploit the hardware layer: GPU architecture, memory hierarchies, CUDA execution model, cuDNN's role, profiling real workloads, and compiling models with TensorRT to hit hard latency SLOs. Learners profile a slow pipeline and ship a measured speedup.
What this course covers.
5 modules, 33 named skill atoms. Expand any module to see them.
1GPU Architecture7 skill atoms
2CUDA Execution Model7 skill atoms
3Profiling Discipline6 skill atoms
4TensorRT Optimization7 skill atoms
5SLO Lab6 skill atoms
AI-209 in the role journeys.
This course appears in 2 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.
Roles this course serves
This course is authored to band B5.
Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.
What do B1–B6 mean?The CASI Capability Ladder — click to expand
Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.
A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.
Other AI courses at this level.
Model Compression: Quantization, PEFT, QLoRA & Distillation
Ray & Large-Scale Model Training
Preference Tuning: RLHF, DPO & Reward Models
Knowledge Graphs & GraphRAG
Multi-Agent Systems & Orchestration
LLM Inference Optimization & Serving (vLLM · LiteLLM)
Run AI-209 for your team.
This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.