← All 186 courses
AI-218 AI & Data

Reinforcement Learning Foundations

Learn reinforcement learning as an engineering discipline rather than a set of demos: frame a problem as an MDP, implement both value-based and policy-gradient agents, and read the failure signatures when learning stalls or the agent games its reward. Learners train agents in Gymnasium environments, diagnose reward hacking they caused themselves, and finish able to argue honestly when bandits, forecasting or plain heuristics beat RL on a business problem.

B4Professional
6modules
42skill atoms
2role journeys
Curriculum

What this course covers.

6 modules, 42 named skill atoms. Expand any module to see them.

1MDPs and the RL Framing7 skill atoms
states actions rewards and transitionsdiscount factor and effective horizonepisodic vs continuing tasksreturn and value definitionsBellman equationspartial observability and POMDPsframing a business problem as an MDP
2Value-Based Methods7 skill atoms
value iteration and dynamic programmingMonte Carlo returnstemporal difference updatesQ-learning off-policy targetSARSA on-policy contrastDQN with function approximationreplay buffer and target network stability
3Policy Gradient Methods7 skill atoms
REINFORCE and the log-derivative trickbaselines for variance reductionactor-critic architecturesadvantage estimation with GAEPPO clipping and trust regionsSAC for continuous action spacesentropy bonus and policy collapse
4Exploration versus Exploitation7 skill atoms
epsilon-greedy annealing schedulesBoltzmann action selectionUCB and Thompson sampling banditsoptimistic initialisationintrinsic curiosity rewardssparse-reward exploration failurecontextual bandits as the cheaper answer
5Reward Design and Its Pathologies7 skill atoms
potential-based reward shaping invariancereward hacking and specification gamingsparse vs dense reward trade-offunintended side effects on ignored variablesGoodhart drift on a proxy rewardsafety constraints and constrained MDPsreward review treated as design work
6Simulation, Sample Efficiency and Fit7 skill atoms
Gymnasium environment APIStable-Baselines3 reference agentssim-to-real gap and domain randomisationsample efficiency versus supervised alternativesoffline RL from logged interaction dataseed variance and honest reportingwhen RL is the wrong tool entirely
Where it fits

AI-218 in the role journeys.

This course appears in 2 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.

AI / ML Engineer

Professional stage
AI-217AI-218AI-510

Data Scientist

Professional stage
AI-217AI-218AI-311

Roles this course serves

The capability ladder

This course is authored to band B4.

Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.

What do B1–B6 mean?The CASI Capability Ladder — click to expand

Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.

What the learner can doTypical evidence
B1
AwareUnderstands concepts and vocabulary; uses tools with guidance
Knowledge checks
B2
FoundationPerforms standard tasks correctly in familiar contexts
Guided labs, autograded exercises
B3
PractitionerDelivers complete pieces of work independently
Scenario labs, proctored hands-on exams
B4
ProfessionalHandles production-grade complexity, trade-offs and failure modes
Break-fix drills, design defenses
B5
AdvancedEngineers systems end-to-end under constraints; leads others
Rubric-scored capstones, vivas
B6
ExpertSets direction; recognised authority across teams
Portfolio + panel evaluation

A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.

Next step

Run AI-218 for your team.

This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.

Add it to a training plan Talk to our team Check your team’s level free