← All 186 courses
AI-311 AI & Data

LLM Evaluation & Benchmarking

Institutionalise 'every claim measured': build golden datasets, automated evaluation harnesses, LLM-as-judge with calibration checks, regression gates in CI, and honest reporting — including designing evals that catch the failure a stakeholder will find first.

B4Professional
5modules
34skill atoms
6role journeys
Curriculum

What this course covers.

5 modules, 34 named skill atoms. Expand any module to see them.

1Eval Strategy6 skill atoms
capability vs safety evalsgolden setstask-representative samplingdataset size & coveragefrozen vs rolling test setscontamination checks
2Metrics that Mean Something7 skill atoms
exact/semantic matchgrounding %rubric scoresembedding similarity thresholdspass@kinter-rater agreementper-slice reporting
3LLM-as-Judge7 skill atoms
judge promptsbias checkshuman calibrationposition & verbosity biaspairwise vs pointwise scoringjudge model choicejudge cost control
4Regression Gates7 skill atoms
eval-in-CIthresholdsrelease blockingper-PR eval subsetsnightly full runsflakiness handlingsign-off record
5Failure-First Design7 skill atoms
adversarial casesreport writingfix loopred-team prompt setsedge-case mining from logsfailure taxonomyfix-and-reverify
Where it fits

AI-311 in the role journeys.

This course appears in 6 of our 45 role journeys. Here is what a learner takes immediately before and after it in each.

GenAI / LLM Engineer

Professional stage
AI-207AI-311AI-405

Data Scientist

Capstone stage
AI-218AI-311AI-510

AI Architect

Practitioner stage
AI-405AI-311AI-310

Forward Deployed Engineer

Professional stage
AI-502AI-311AI-501

AI Quality / Eval Engineer

Practitioner stage
AI-101AI-311AI-310

QA / Test Engineer (AI-era)

Professional stage
DEV-302AI-311CL-403

Roles this course serves

The capability ladder

This course is authored to band B4.

Every course we run is written to one rung of the CASI ladder, so a plan can be assembled to take a team from where they are to where they need to be.

What do B1–B6 mean?The CASI Capability Ladder — click to expand

Every course targets a band on the CASI Capability Ladder — our six-band proficiency scale, anchored to open standards (O*NET, ESCO, NICE, NIST AI RMF, Bloom's). A band tells you how deep a course goes, and what evidence proves it.

What the learner can doTypical evidence
B1
AwareUnderstands concepts and vocabulary; uses tools with guidance
Knowledge checks
B2
FoundationPerforms standard tasks correctly in familiar contexts
Guided labs, autograded exercises
B3
PractitionerDelivers complete pieces of work independently
Scenario labs, proctored hands-on exams
B4
ProfessionalHandles production-grade complexity, trade-offs and failure modes
Break-fix drills, design defenses
B5
AdvancedEngineers systems end-to-end under constraints; leads others
Rubric-scored capstones, vivas
B6
ExpertSets direction; recognised authority across teams
Portfolio + panel evaluation

A note on B6. Courses in this catalog target B1–B5. B6 is not taught — it is recognised, through a portfolio and a panel, once someone is setting direction for others. Every journey here is built to land a learner at B5.

Next step

Run AI-311 for your team.

This course runs at several lengths depending on how deep you need to go and how much of it your people already have. Tell us who is being trained and we will scope it.

Add it to a training plan Talk to our team Check your team’s level free