Evaluation
You judged the output. You noticed what failed and named it.
Judge output// 00 Loop engineering · the masterclass, made interactive
Most people still use AI one prompt at a time. Loop engineering takes the four jobs you do in your head — evaluation, memory, guardrails, stopping conditions — and builds them into a system that can tell whether its own work is improving.
You design the controls. The loop runs them. The one-page spec keeps both honest — failures flow back to you as evidence, not as vibes.
Real work has a quality bar, happens at volume, and needs to stay consistent long after your attention starts to drift. At ten requests you read carefully. At request fifty you start scanning. By request two hundred, the model still generates at the same speed — your quality control does not.
You judged the output. You noticed what failed and named it.
Judge outputYou remembered what had failed and carried it into the next attempt.
Carry lessonsYou enforced the boundaries — no invented delivery dates, no broken tone.
Hold boundariesYou decided when the answer was ready. The model never could.
Declare done→ Without naming any of it, you were running a loop manually, inside your head.
Drafts one reply. If the reply is weak, nothing happens unless a human notices.
Retrieves, drafts, routes — a sequence designed beforehand.
Receives an objective and tools; asks for clarification when sources conflict.
Evaluates against a defined standard. Failures become input to the next attempt. Passing checks stop the run; conflicting evidence escalates.
Treat a loop like a new employee with infinite patience and no earned judgment.AI evaluates the fiftieth item with the same energy as the first.
The model handles generation and reasoning. The system around it supplies the controls. Select any part to see the matching lines of the loop light up.
A measurable state, not an activity
“Review today's pull requests” describes an activity without defining when the review is complete. Inspect every changed file, run the required tests, link failures to evidence, flag security-sensitive changes, recommend approve / request-changes / escalate.
If the goal is an activity, the loop has no definition of done.
while target_not_reached:result = take_action(context, memory)evaluation = evaluate(result, rubric)log(result, evaluation)if evaluation.passes:stop_with_success(result)if human_decision_required(evaluation):escalate(result, evaluation)if budget_exhausted() or progress_stalled():stop_without_success()memory = update_memory(result, evaluation)
Every proposed change is a challenger. A challenger takes the title only if it proves it is better — on a holdout it never studied while being proposed.
Propose → test → compare → promote or reject → log → repeat.
The current version holds the title until a challenger beats it.
Evidence — 7 rounds · $1.94 · holdout 3.4 → 4.6.
Batches of 20; every cluster keeps an original quote and source.
Gate — a human traces every claim to a customer's words.
One model builds; another argues it is wrong. Objections are Fixed or Accepted on record.
Evidence — 11 objections → 7 changes · 4 accepted deliberately.
Fresh browser session, real task, score the whole journey, fix the weakest screen, restart.
Evidence — "Invalid input." → named field, format, fix. Screen 2 → 4.
Shorter replies passed more checks. By round nine the answers were brief, polite, safe — and almost useless.
If the judge is easier to fool than the task, the loop learns to please the judge.
One request becomes many rounds and passes. Rejected challengers still burned tokens.
A loop that works only while someone watches its terminal is still a prototype.
Perfection unreachable → the loop spins and the budget bleeds.
Runs on real work, affects nothing. Compare with the human decision.
BoundaryRead-only. Output: a shadow decision, logged next to the human one.
If you cannot explain how you would judge it yourself, clarify good work before automating it.
Automate what is checkable. Escalate what is consequential.
"We will see how it goes" gives an unattended system no operating policy.
The blank shows which part of the system still needs design work.
One change preserves causality.
Improvement-set wins are rumors until they survive the holdout.
A tie goes to the champion.
These need a human gate, not a better prompt.
Shadow → suggest → approve → bounded, on observable evidence.
Fill this in before building — the document on the right updates live on every click and keystroke.