Get your team genuinely good at OpenAI and Claude
Two structured tracks, five levels each, from "what is a token" to running evaluated, governed agents in production. Every lesson carries a hands-on exercise and a knowledge check, and your progress is saved as you go.
The two tracks
Each track stands alone โ you do not need one to follow the other. Most teams send engineers through both and everyone else through Level 1 and Level 2 of the platform they actually use.
Claude
The Claude model family, prompting for current-generation models, the Messages API, tool use and MCP, Managed Agents, and production governance.
- L1 โ Foundations: models, pricing, surfaces
- L2 โ Prompting: structure, thinking, effort, unlearning
- L3 โ Building: Messages API, caching, structured outputs
- L4 โ Tools and agents: tool use, server tools, MCP, design
- L5 โ Expert: Managed Agents, memory, evals, migration
OpenAI
The GPT-5.6 tiers, prompting and reasoning effort, the Responses API, structured outputs and RAG, the Agents SDK, and reliability engineering.
- L1 โ Foundations: tiers, surfaces, the cost model
- L2 โ Prompting: specificity, effort, systematic debugging
- L3 โ Building: Responses API, function calling, retrieval
- L4 โ Agents: built-in tools, Agents SDK, MCP
- L5 โ Expert: evals, reliability, governance
Where to start
Four paths through the same material, depending on who you are.
Non-technical โ 90 minutes
Level 1 and Level 2 of the platform your team already uses, then the Glossary. You will finish able to write prompts that work, choose the right tool for a task, and explain to a colleague why the model made something up.
Engineer, new to LLMs โ one day
Levels 1 to 3 of one track end to end, doing every exercise. Then Labs 1 through 4. You will finish with a working, instrumented integration rather than a notebook.
Engineer, already shipping โ half a day
Skim Levels 1 and 2, then do Levels 4 and 5 properly. Level 5 is where most production problems actually live: evals, migration, cost, and governance.
Tech lead or manager โ two hours
Level 1 of both tracks for the comparison, then Level 5 of both for evaluation, reliability, and governance. Then run the Labs section as a team workshop.
Running this with your team
The content is the easy part. This is what makes it stick.
1. Cohorts, not self-service
Six to ten people moving through one level a week, with a 45-minute session at the end of each to compare exercise results. Self-paced courses have a completion rate that rounds to zero; cohorts do not.
2. Real work, not toy examples
Every exercise is written to be done against your own data and your own systems. Substitute a real task at every opportunity โ the learning transfers, the toy examples do not.
3. Ship one thing per level
A shared prompt library after Level 2. A working integration after Level 3. An eval set after Level 5. Training that produces artifacts survives contact with the next quarter's priorities.
About your progress
Progress is stored in this browser only โ nothing is sent anywhere, and there is no account. That means it does not follow you between devices, and clearing site data clears it. For team reporting, track completion however you already track training.