courses / course 1 · the agent engineering program
Agent School
Learn to build agents that are actually useful.
A guided program run inside Claude Code. You design and ship real agents. Our harness verifies what they can do. The proof lives in your own GitHub.
The premise
Models change every month. Prices drop overnight, new categories appear, and yesterday's best choice becomes today's expensive habit. So this program does not teach you a model. It teaches you the method: measure before you trust, version everything, design for swap, and prove your system still works when the ground moves. Most of what we teach about specific models will be obsolete within a year. That is the point. You will be the person who is fine with that.
How it works
- 1
Missions, not lectures. You get a real situation, constraints and a budget. You decide, your agent builds, reality responds. The principle arrives in the debrief, after you have lived it.
- 2
Your repos, your name. Every project ships to your own GitHub with an eval suite, a cost log and your decision record. Nothing is branded coursework.
- 3
Verified, not certified. Our harness drives your agents against hidden test sets and adversarial inputs. The numbers it measures are the credential, and they are reproducible by anyone.
- 4
Runs on your own subscription. Your agents run on the Claude Pro you already have. Budgets are still part of the curriculum, because real systems have bills, and the harness measures cost itself when it verifies your work.
- 5
The ground moves on purpose. Mid-program, a model you depend on gets deprecated or doubles in price. Your evals tell you what to do next. That is the taste of real life.
The path
| Stage | You build | It proves |
|---|---|---|
| Pre-work | Foundations: your tools, your first conversation with a model, and the field notes that explain what is actually happening | you understand tokens, context and cost before anything depends on them |
| Warm-up | A job-watch agent on our live MCP server | you can ship a working agent in an evening |
| Project 1 | A research and enrichment agent for a vertical you choose | structured output, verification, cost control |
| Project 2 | A document pipeline on real, messy data, graded against hidden labels | ground truth, precision and recall, idempotent reruns |
| Project 3 | A knowledge agent: question answering over the real documents of your vertical | retrieval quality you can measure, graded on golden questions |
| Project 4 | An agent that acts in a real workflow, with approval gates and an audit log | safety, consent, failure handling |
| Capstone | One of your agents packaged as a product: an MCP server with docs, auth and a published eval report | you can ship something other systems rely on |
Every concept, covered
If you are comparing syllabi: the concepts on a standard forward-deployed-engineer curriculum are all here. The difference is how you meet them. Each one arrives as a short field note at the moment a project makes it matter, never as a lecture to sit through.
| Concept | Where you learn it by building |
|---|---|
| AI deployment mindset and FDE responsibilities | Pre-work framing, then every debrief: you run the whole loop from brief to verified result |
| Business problem discovery and requirement mapping | Projects 1 and 4: you interview a real stakeholder and write the requirement one-pager before anything gets built |
| GenAI and LLM foundations | Pre-work and week 1: you watch tokens get billed and context fill up inside your own first agent |
| Prompt engineering for business use cases | Projects 1 and 2: structured outputs, schemas and examples, graded by evals rather than vibes |
| RAG systems and enterprise knowledge workflows | Project 3: retrieval over real documents, graded on golden questions, citations required |
| Agentic AI workflows and automation design | Project 4: an agent that acts, with approval gates and an audit log |
| APIs, data pipelines and system integration | Every project moves real data; the capstone publishes your own API as an MCP server |
| AI governance, privacy, risk and guardrails | Approval gates, consent and audit logs in Project 4, plus the field notes on what goes wrong in production |
| The anatomy of an agent: state, skills, memory, multi-agent design | Cluster 7 of the field notes, exercised from Project 1 through the capstone |
| Self-improving systems and decision traces | Cluster 10: the improvement loop, ADRs, and gated self-improvement in your capstone |
| Stakeholder communication and solution presentation | A one-page solution memo per project, and a recorded walkthrough of your capstone |
What you will build
Every project ships to your own GitHub, in a vertical you choose. These are examples of what past-you could not build and future-you will.
1
A research and enrichment agent
Pick a vertical you know: real-estate brokers in Pune, D2C skincare brands, NGO grant pipelines. Your agent takes a list of 300 targets and returns verified, structured fields with cited sources, under a budget you set.
Example output
300 companies, 12 fields each, 94% field accuracy against your own 50-row hand check, $4.80 total inference, every claim linked to its source.
2
A document pipeline
Messy real documents in, clean validated records out. Resumes for a placement cell, invoices for a CA firm, or our own corpus of live job postings, where your results are graded against labels you never see.
Example output
1,000 documents processed, precision and recall reported per field, identical database state across three reruns, cost per document in the readme.
3
A knowledge agent
Retrieval-augmented answers over a real document set from your vertical: a society’s bylaws and minutes, a clinic’s SOPs, a firm’s policy folder. You build the golden-question set first, then make the agent pass it, with every answer citing its source.
Example output
60 golden questions, answer accuracy and citation accuracy reported separately, refuses to answer when the documents do not contain the answer.
4
An agent that acts
Triage for a real inbox or queue, with a human approval gate on anything irreversible and an audit log of every action. You will find one real user: a club, a shop, a team at work. That one real user changes everything about how you build.
Example output
Support triage for a small Shopify store, 80 messages a week, drafted replies approved by the owner, zero unapproved sends, full audit trail.
5
Your agent as a product
The capstone packages one of your agents as an MCP server other agents can call: authentication, rate limits, documentation, and a published eval report. Our harness consumes it cold, with no help from you. If it works, that is the proof.
Example output
A live MCP endpoint, a third-party agent completing a task against it unaided, and an eval report page you can link in any application.
The eight weeks
Pre-work unlocks the day you enroll, and missions unlock by progress, not by calendar, so nobody is left behind. The cohort moves roughly like this.
| Week | You are doing | It ends with |
|---|---|---|
| Pre | Pre-work: install your tools, meet your agent, read the foundations notes, and build the warm-up job-watch agent | your tools work, the vocabulary is yours, and your first agent is in your repo |
| 1 | Project 1 begins: find the problem, interview one real person, write the page your agent will execute | a one-pager your agent can execute, and your first cost estimate on record |
| 2 | Ship and verify Project 1: a research agent with verified numbers, not vibes | verified accuracy and cost figures in your repo, checked against your own hand-labeled sample |
| 3 | Project 2 begins: ground truth first. You build the exam before you build the student | a labeled set, an eval written by you, and a pipeline taking shape against it |
| 4 | The ladder lab, and shipping against hidden labels: the heaviest week, deliberately | precision, recall and cost per document verified by the harness, and a routing decision you can defend |
| 5 | Project 3: the knowledge agent. Golden questions first, and midweek the ground moves | a golden-question set, a retrieval design, and a migration record showing you survived the change |
| 6 | Ship Project 3: answer accuracy, citation accuracy and refusal correctness, never merged | measured answer and citation accuracy over documents you chose, and a refusal path you designed |
| 7 | Project 4: your agent acts in the real world, for a real user, behind approval gates | a deployed agent with an audit log, and its first real traffic beginning to accrue |
| 8 | Capstone: package, publish, present. One of your agents becomes a product other systems can rely on | a live MCP endpoint, a published eval report, your decision log, and the final debrief |
A week in the program
Monday, the mission brief drops and you write a one-page prediction sheet: expected cost, expected failures, expected score. Thirty minutes that change how you build. Through the week you work with your agent in Claude Code, on your schedule; most weeks fit in 8 to 12 hours. The harness runs whenever you are ready, as many times as you like. Friday is the live debrief: the cohort's predictions against the cohort's results, the failure gallery, and the one principle the week actually taught. Office hours sit midweek for whoever is stuck.
What the harness verifies (and two labeled examples)
The harness is a grading system that drives your agent the way reality would. It runs your pipeline against inputs you have never seen, including documents designed to fool it. It re-runs your system twice and diffs the state. It kills your orchestrator mid-run and checks that resuming does not corrupt anything. It meters every token your agents spend and signs the numbers. What it cannot do is be impressed by a readme.
eligibility screener, verified 2026-11-xx precision 97.4% / recall 96.1% (n=500, hidden set) cost per document: $0.0037, gateway-metered idempotency: identical state across 3 reruns survived: week-5 provider price shock (see ADR-3)
ADR-4: moved seniority classification from the frontier model to a small one. Accuracy held within 1.1 points on our eval; cost fell 8x. Accepted the trade. Exec-level titles now route to the escalation rule instead.
What it costs, all in
| Item | Student |
|---|---|
| Program fee (incl. GST) | ₹9,990 |
| Claude Pro, 2 months | about ₹3,800, you keep it after |
| All-in | about ₹13,800 |
No other costs. Your agents run on your own subscription, and verified cost figures are measured by our harness when it grades your work.
Compare the all-in number to any program that quotes a fee and stays quiet about the AI bill.
The line on your resume
Each project is written to earn one specific sentence. Project 1: built a research agent that enriched 300 accounts at 94% field accuracy for under five dollars. Project 2: shipped a document pipeline graded at 97% precision against a hidden test set. Project 3: built a documentation Q&A agent with measured citation accuracy. Project 4: deployed a human-in-the-loop triage agent for a real business. Capstone: published an MCP server with a public eval report. The figures on your resume will be yours, not these examples, and anyone who doubts them can run your evals.
Who runs this
The program is run by the team behind Deployed, the job board you are on right now. The classification pipeline behind this site has processed over 139,000 live postings, and the harness that will grade your work applies the same discipline we use in production: measure before trusting, version everything, survive the model churn. The missions are distilled from building this exact product.
What you leave with
Four or five working agents in your own repositories, each with measured accuracy and cost figures our harness verified. A decision log that shows how you think. A solution memo per project, written for a reader who does not code. A profile on Deployed where employers see the verified numbers next to your name. And the habit that outlasts every model release: measure, version, verify.
Live, but not lectures
Weekly live debriefs review the cohort's own runs: where predictions missed, what the failures have in common, who made the call that worked. There is no lecture library to fall behind on. Reality lectures. We debrief.
Field notes, not textbooks
The concepts live in forty-two field notes: plain-language explainers, ten minutes each, one concept per note. Each note opens with a question, explains the idea with one analogy, tells one true story from our own production systems, and ends with a two-minute experiment you run yourself. They are sequenced so you can read them cover to cover like a short book, and they sit as files inside your course repo, which means your agent can re-explain any of them using your own project as the example. Ask it to try a different analogy, or to quiz you. It will not get tired.
Pricing
All prices include GST.
Student cohort
₹9,990
The full path. Weekly live debriefs. You bring your own Claude Pro subscription (about ₹1,900 per month, the same tool you will use at work). Full refund through the end of week one.
Status: Cohort 1 · waitlist open
Professional
₹24,999
Everything in Student, plus Claude Pro included for two months, mentor circles with practicing forward-deployed engineers, and fast-track consideration for Stak's placement bench.
Status: Opens with cohort 1
Self-paced
₹6,999
Same missions and harness, agent mentor only, start any time.
Status: Later this year
Who this is for
New to AI is the expected starting point: the pre-work and the field notes assume you have never touched a model. What you do need is basic programming literacy. You can read and write simple code in some language, even if an agent writes most of yours here. Students and early-career engineers fit the student tier. Working engineers aiming at forward-deployed and AI engineering roles fit the professional tier.
FAQ
Are there live lectures?
Live debriefs every week, reviewing the cohort's own work. No lecture library.
Do I need to be a strong coder?
You need basic programming literacy: you can read simple code and reason about what it does. Your agent writes most of the code. You decide what gets built and prove whether it works.
I know nothing about AI. Can I keep up?
Yes, that is the audience this is designed for. The pre-work covers the foundations before anything depends on them, every concept has a plain-language field note, and your agent can re-explain any note with your own project as the example.
Is there material I can read and study?
Forty-two field notes: short plain-language explainers covering every concept from LLM foundations to RAG to governance. Readable cover to cover in a weekend. The free ones are readable now at /courses/notes, and they live in your repo so your agent can expand on any of them.
What tools do I need?
Claude Code (or Codex) and a Claude Pro subscription (about ₹1,900 per month). Your agents run on that same subscription. Everything else is included.
Do you guarantee a job?
No. We verify what you built and make it visible to employers on Deployed. Stak may present top graduates to hiring companies, with your consent. Nobody serious guarantees placements.
When does it start?
Cohort 1 is being assembled now. Join the waitlist and we will email you the dates.
Refunds?
Full refund through the end of week one of your cohort, no questions.
How much time does it take?
Plan for 8 to 12 hours a week for 8 weeks, plus light pre-work you can finish before the cohort starts. Missions are sized for evenings and one weekend block.
What if I fall behind?
Missions unlock by your progress, not the calendar. Debrief recordings cover what you miss, and you keep access for six months after your cohort.
I can code but I have never built with LLMs.
That is the expected starting point. The warm-up exists so your first working agent happens in week one.
Can I use Codex instead of Claude Code?
Yes. Missions and the harness are tool-agnostic. Debriefs demo in Claude Code.
Do I get a certificate?
No. You get verified numbers in your own repos and on your Deployed profile. Anyone can check them, which is the point. Certificates ask to be trusted; evals ask to be run.
Can my company buy seats?
Yes. Join the waitlist and mention it in the note field.
Read more
Join the waitlist
Operated by Stak. Your email is used only to contact you about this program.