deployed_

the weeks / Week 4 · Project 2

The ladder lab, and shipping against hidden labels

The heaviest week, deliberately.

Run the same documents through models from cheapest to frontier and measure, on your own held-out set, what each step of the ladder is worth. Then ship: your pipeline graded against labels you have never seen, rerun twice with the database diffed, killed mid-run and resumed. Your harness ships with the pipeline. You also write your first solution memo: one page, for a reader who does not code.

It ends with

Precision, recall and cost per document verified by the harness, a routing decision you can defend with measurements, and the memo.

Field notes

  • 6.4 What happens when you run it twice? program
  • 6.5 Why does your agent need a harness of its own? program
  • 9.2 How should a system fail? program
  • 10.1 Why doesn't the model learn from its mistakes? program
  • 10.2 What is a decision trace, and why write down the why? program
  • 11.1 How do you explain your system to someone who will never read the code? program