deployed_

courses / course 2

RL Environments & Evals

The work AI labs post as RL Environment Engineer and evals roles: authoring verifier-backed tasks, auditing LLM judges, QA-ing agent trajectories.

This is the RL environments and evals work AI labs and training-data companies pay for by the hour, from the harness layer around models to the tasks that grade them. The course is built backward from those real engagements, so every sprint produces something a lab or data vendor would recognise. You finish with a published, verifier-backed task pack as your work sample.

Prerequisite: Agent School (Course 1) or equivalent experience building and evaluating agents.

How the course runs

Six sprints. Every sprint runs the same loop.

  1. 1

    Read with a question. You go into real material with something specific to find out, not to skim.

  2. 2

    Produce an artifact. What you learned has to become something you made.

  3. 3

    Defend it orally. An oral defense, graded, with evidence required for every claim you make.

The skills ladder, one rung per sprint

  1. 1Reading real grading stacks
  2. 2The market's task format
  3. 3Judge auditing
  4. 4Task authoring at difficulty
  5. 5Trajectory QA
  6. 6Shipping a work sample

Open reading list

The course's core sources are public. The course adds the guided walkthrough, the graded defenses, and the task gym.

The doctrine: what buyers of agent tasks actually pay for

Join the waitlist

Status: Cohort 1 — waitlist open.