Transformers from First Principles
Open the black box: self-attention, multi-head attention, positional encoding, encoder/decoder stacks — implemented small enough to read, then mapped onto production models. Learners can afterwards explain, diagram and defend every arrow in the transformer architecture.
What this course covers.
5 modules, 30 named skill atoms. Expand any module to see them.
1Attention Mechanics7 skill atoms
2Full Architecture7 skill atoms
3Build a Mini-Transformer5 skill atoms
4From Paper to Production6 skill atoms
5Whiteboard Certification5 skill atoms
AI-206 in a role journey.
This course appears in 1 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.
Roles this course serves
This course is authored to band B4.
Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.
What do B1–B6 mean?The CASI Capability Ladder — click to expand
Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.
A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.
Other AI courses at this level.
Departmental AI Adoption & Automation Design
Fine-Tuning with Hugging Face
Distributed Training Foundations
Synthetic Data Generation & Evaluation
Vector Databases & Hybrid Search
Document AI & Intelligent Document Processing
Run AI-206 for your team.
This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.