courses / course 2
RL Environments & Evals
The work AI labs post as RL Environment Engineer and evals roles: authoring verifier-backed tasks, auditing LLM judges, QA-ing agent trajectories.
This is the RL environments and evals work AI labs and training-data companies pay for by the hour, from the harness layer around models to the tasks that grade them. The course is built backward from those real engagements, so every sprint produces something a lab or data vendor would recognise. You finish with a published, verifier-backed task pack as your work sample.
Prerequisite: Agent School (Course 1) or equivalent experience building and evaluating agents.
How the course runs
Six sprints. Every sprint runs the same loop.
- 1
Read with a question. You go into real material with something specific to find out, not to skim.
- 2
Produce an artifact. What you learned has to become something you made.
- 3
Defend it orally. An oral defense, graded, with evidence required for every claim you make.
The skills ladder, one rung per sprint
- 1Reading real grading stacks
- 2The market's task format
- 3Judge auditing
- 4Task authoring at difficulty
- 5Trajectory QA
- 6Shipping a work sample
Open reading list
The course's core sources are public. The course adds the guided walkthrough, the graded defenses, and the task gym.
- The doctrine: what buyers of agent tasks actually pay for
- The format: how the market packages tasks
- Judge auditing: how graders fail
- Difficulty and anti-gaming
Join the waitlist
Status: Cohort 1 — waitlist open.