deployed_

agent engineering program

Learn to build agents that are actually useful.

A guided program run inside Claude Code. You design and ship real agents. Our harness verifies what they can do. The proof lives in your own GitHub.

The premise

Models change every month. Prices drop overnight, new categories appear, and yesterday's best choice becomes today's expensive habit. So this program does not teach you a model. It teaches you the method: measure before you trust, version everything, design for swap, and prove your system still works when the ground moves. Most of what we teach about specific models will be obsolete within a year. That is the point. You will be the person who is fine with that.

How it works

  1. 1

    Missions, not lectures. You get a real situation, constraints and a budget. You decide, your agent builds, reality responds. The principle arrives in the debrief, after you have lived it.

  2. 2

    Your repos, your name. Every project ships to your own GitHub with an eval suite, a cost log and your decision record. Nothing is branded coursework.

  3. 3

    Verified, not certified. Our harness drives your agents against hidden test sets and adversarial inputs. The numbers it measures are the credential, and they are reproducible by anyone.

  4. 4

    Tokens included. Your agents' inference runs on a metered budget we provide. The budget is part of the curriculum, because real systems have bills.

  5. 5

    The ground moves on purpose. Mid-program, a model you depend on gets deprecated or doubles in price. Your evals tell you what to do next. That is the taste of real life.

The path

StageYou buildIt proves
Warm-upA job-watch agent on our live MCP serveryou can ship a working agent in an evening
Project 1A research and enrichment agent for a vertical you choosestructured output, verification, cost control
Project 2A document pipeline on real, messy data, graded against hidden labelsground truth, precision and recall, idempotent reruns
Project 3An agent that acts in a real workflow, with approval gates and an audit logsafety, consent, failure handling
CapstoneOne of your agents packaged as a product: an MCP server with docs, auth and a published eval reportyou can ship something other systems rely on

What you will build

Every project ships to your own GitHub, in a vertical you choose. These are examples of what past-you could not build and future-you will.

1

A research and enrichment agent

Pick a vertical you know: real-estate brokers in Pune, D2C skincare brands, NGO grant pipelines. Your agent takes a list of 300 targets and returns verified, structured fields with cited sources, under a budget you set.

Example output

300 companies, 12 fields each, 94% field accuracy against your own 50-row hand check, $4.80 total inference, every claim linked to its source.

2

A document pipeline

Messy real documents in, clean validated records out. Resumes for a placement cell, invoices for a CA firm, or our own corpus of live job postings, where your results are graded against labels you never see.

Example output

1,000 documents processed, precision and recall reported per field, identical database state across three reruns, cost per document in the readme.

3

An agent that acts

Triage for a real inbox or queue, with a human approval gate on anything irreversible and an audit log of every action. You will find one real user: a club, a shop, a team at work. That one real user changes everything about how you build.

Example output

Support triage for a small Shopify store, 80 messages a week, drafted replies approved by the owner, zero unapproved sends, full audit trail.

4

Your agent as a product

The capstone packages one of your agents as an MCP server other agents can call: authentication, rate limits, documentation, and a published eval report. Our harness consumes it cold, with no help from you. If it works, that is the proof.

Example output

A live MCP endpoint, a third-party agent completing a task against it unaided, and an eval report page you can link in any application.

The eight weeks

Missions unlock by progress, not by calendar, so nobody is left behind. The cohort moves roughly like this.

WeekYou are doingIt ends with
1Setup and the warm-up: a job-watch agent on our live MCP serveryour first harness-verified agent, built in an evening
2Project 1 begins: pick your vertical, write your first prediction sheet, design the budgeta brief your agent can execute, and your first cost estimate on record
3Ship and verify Project 1verified accuracy and cost figures in your repo, and your first live debrief of the cohort's numbers
4Project 2 begins: ground truth. You build a labeled set before you build the pipelinean eval your pipeline must pass, written by you
5The ladder lab, and the ground moves: the same documents run through four models from cheapest to frontier, and mid-week a model you depend on changes price or disappearsa routing decision you can defend with measurements, and a migration record showing you survived the change
6Ship Project 2 against hidden labelsprecision, recall and cost per document, verified by the harness
7Project 3: your agent acts in a real workflow, for a real user, behind approval gatesa deployed agent with an audit log and its first real week of traffic
8Capstone: package, publish, presentan MCP server with a published eval report, your decision log, and the final debrief

A week in the program

Monday, the mission brief drops and you write a one-page prediction sheet: expected cost, expected failures, expected score. Thirty minutes that change how you build. Through the week you work with your agent in Claude Code, on your schedule; most weeks fit in 8 to 12 hours. The harness runs whenever you are ready, as many times as you like. Friday is the live debrief: the cohort's predictions against the cohort's results, the failure gallery, and the one principle the week actually taught. Office hours sit midweek for whoever is stuck.

What the harness verifies (and two labeled examples)

The harness is a grading system that drives your agent the way reality would. It runs your pipeline against inputs you have never seen, including documents designed to fool it. It re-runs your system twice and diffs the state. It kills your orchestrator mid-run and checks that resuming does not corrupt anything. It meters every token your agents spend and signs the numbers. What it cannot do is be impressed by a readme.

Example: the metrics block your repo will carry
eligibility screener, verified 2026-11-xx
precision 97.4% / recall 96.1% (n=500, hidden set)
cost per document: $0.0037, gateway-metered
idempotency: identical state across 3 reruns
survived: week-5 provider price shock (see ADR-3)
Example: a decision record from a past build

ADR-4: moved seniority classification from the frontier model to a small one. Accuracy held within 1.1 points on our eval; cost fell 8x. Accepted the trade. Exec-level titles now route to the escalation rule instead.

What it costs, all in

ItemFoundingStudent
Program fee (incl. GST and your agents' token budget)₹5,999₹9,990
Claude Pro, 2 monthsincludedabout ₹3,800, you keep it after
All-in₹5,999about ₹13,800

No other costs. Token top-ups exist if you burn past the included budget; most learners will not.

Compare the all-in number to any program that quotes a fee and stays quiet about the AI bill.

The line on your resume

Each project is written to earn one specific sentence. Project 1: built a research agent that enriched 300 accounts at 94% field accuracy for under five dollars. Project 2: shipped a document pipeline graded at 97% precision against a hidden test set. Project 3: deployed a human-in-the-loop triage agent for a real business. Capstone: published an MCP server with a public eval report. The figures on your resume will be yours, not these examples, and anyone who doubts them can run your evals.

Who runs this

The program is run by the team behind Deployed, the job board you are on right now. The classification pipeline behind this site has processed over 139,000 live postings, and the harness that will grade your work applies the same discipline we use in production: measure before trusting, version everything, survive the model churn. The missions are distilled from building this exact product.

What you leave with

Three or four working agents in your own repositories, each with measured accuracy and cost figures our harness verified. A decision log that shows how you think. A profile on Deployed where employers see the verified numbers next to your name. And the habit that outlasts every model release: measure, version, verify.

Live, but not lectures

Weekly live debriefs review the cohort's own runs: where predictions missed, what the failures have in common, who made the call that worked. There is no lecture library to fall behind on. Reality lectures. We debrief.

Pricing

All prices include GST and the token allowance.

Founding cohort

₹5,999

25 seats. Warm-up + Project 2, personal review, Claude Pro included for the cohort, full refund through week one.

Status: Waitlist open

Student cohort

₹9,990

The full path. Weekly live debriefs. You bring your own Claude Pro subscription (about ₹1,900 per month, the same tool you will use at work).

Status: After the founding cohort

Professional

₹24,999

Everything in Student, plus Claude Pro included for two months, mentor circles with practicing forward-deployed engineers, and fast-track consideration for Stak's placement bench.

Status: After the founding cohort

Self-paced

₹6,999

Same missions and harness, agent mentor only, start any time.

Status: Later this year

Top-ups for the token allowance are available if you burn past it. Most learners will not.

Who this is for

You can read code, even if an agent writes most of yours. You want proof of skill that survives scrutiny, not a certificate. Students and early-career engineers fit the student tier. Working engineers aiming at forward-deployed and AI engineering roles fit the professional tier.

FAQ

Are there live lectures?

Live debriefs every week, reviewing the cohort's own work. No lecture library.

Do I need to be a strong coder?

You need to read code and reason about systems. Your agent writes most of the code. You decide what gets built and prove whether it works.

What tools do I need?

Claude Code (or Codex) and, outside the founding cohort, a Claude Pro subscription. Everything else is included.

Do you guarantee a job?

No. We verify what you built and make it visible to employers on Deployed. Stak may present top graduates to hiring companies, with your consent. Nobody serious guarantees placements.

When does it start?

The founding cohort is being assembled now. Join the waitlist and we will email you the dates.

Refunds?

Founding cohort: full refund through the end of week one, no questions.

How much time does it take?

Plan for 8 to 12 hours a week for 8 weeks. Missions are sized for evenings and one weekend block.

What if I fall behind?

Missions unlock by your progress, not the calendar. Debrief recordings cover what you miss, and you keep access for six months after your cohort.

I can code but I have never built with LLMs.

That is the expected starting point. The warm-up exists so your first working agent happens in week one.

Can I use Codex instead of Claude Code?

Yes. Missions and the harness are tool-agnostic. Debriefs demo in Claude Code.

Do I get a certificate?

No. You get verified numbers in your own repos and on your Deployed profile. Anyone can check them, which is the point. Certificates ask to be trusted; evals ask to be run.

Can my company buy seats?

Yes. Join the waitlist and mention it in the note field.

Join the waitlist

Operated by Stak. Your email is used only to contact you about this program.