// 00 Loop engineering · the masterclass, made interactive

Act → evaluate → improve, live.

Most people still use AI one prompt at a time. Loop engineering takes the four jobs you do in your head — evaluation, memory, guardrails, stopping conditions — and builds them into a system that can tell whether its own work is improving.

Source · Babete - Loop Engineering ResearchTheme engine · data-theme swaps tokens, not markup9 parts · 4 architectures · 5 failure modes
ACTTake one narrow action with current context and memory.
⟲ Repeat■ Stop↑ Escalate
// 00.5Orbit viewFig. 0 — how every loop works

How every loop works.

You design the controls. The loop runs them. The one-page spec keeps both honest — failures flow back to you as evidence, not as vibes.

The shared contract between a person and an autonomous loopThe person specifies, grants, audits, and shadows. The loop acts, observes, evaluates, updates, and decides. A shared specification controls triggering and escalation.Youyou design the controlsThe Loopit runs itselfLoop SpecSHARED CONTRACTgoal · rubric · guardrails · stopSPEC INEVIDENCE OUTSpecifyone-page spec · one change per roundGrantautonomy on evidence · sample itTrigger runspec in · budget · guardrails onActone narrow action · toolsObserveresult + evidenceEvaluaterubric · concrete checksUpdatestate object · memoryDeciderepeat · stop · escalateEscalationconflict · evidence · decision requiredAuditread logs · strengthen the evaluatorShadowreal work · zero authority↩ failures become input to the next attempt
// 01The problemFig. 1

One prompt eventually stops working.

Real work has a quality bar, happens at volume, and needs to stay consistent long after your attention starts to drift. At ten requests you read carefully. At request fifty you start scanning. By request two hundred, the model still generates at the same speed — your quality control does not.

Job / 01

Evaluation

You judged the output. You noticed what failed and named it.

Judge output
Job / 02

Memory

You remembered what had failed and carried it into the next attempt.

Carry lessons
Job / 03

Guardrails

You enforced the boundaries — no invented delivery dates, no broken tone.

Hold boundaries
Job / 04

Stopping

You decided when the answer was ready. The model never could.

Declare done

→ Without naming any of it, you were running a loop manually, inside your head.

Prompt

Drafts one reply. If the reply is weak, nothing happens unless a human notices.

Workflow

Retrieves, drafts, routes — a sequence designed beforehand.

Agent

Receives an objective and tools; asks for clarification when sources conflict.

Loop

Evaluates against a defined standard. Failures become input to the next attempt. Passing checks stop the run; conflicting evidence escalates.

Treat a loop like a new employee with infinite patience and no earned judgment.AI evaluates the fiftieth item with the same energy as the first.
// 02AnatomyFig. 2 — spec inspector

The nine parts of a reliable loop.

The model handles generation and reasoning. The system around it supplies the controls. Select any part to see the matching lines of the loop light up.

01 — Goal

A measurable state, not an activity

“Review today's pull requests” describes an activity without defining when the review is complete. Inspect every changed file, run the required tests, link failures to evidence, flag security-sensitive changes, recommend approve / request-changes / escalate.

If the goal is an activity, the loop has no definition of done.
loop.py · The architectureAct → Observe → Evaluate → Update → Decide
while target_not_reached:
    result     = take_action(context, memory)
    evaluation = evaluate(result, rubric)
    log(result, evaluation)
    if evaluation.passes:
        stop_with_success(result)
    if human_decision_required(evaluation):
        escalate(result, evaluation)
    if budget_exhausted() or progress_stalled():
        stop_without_success()
    memory = update_memory(result, evaluation)
// 03Live runFig. 3 — champion loop

The champion defends its title.

Every proposed change is a challenger. A challenger takes the title only if it proves it is better — on a holdout it never studied while being proposed.

State: Idle · awaiting run
012345Target 4.5BaseR1R2R3R4R5R6R7
—Champion
—/12Rounds
0/3Stall
0 · 0Promo · Rej
0 minUnattended
$0.00Cost
Standby

Propose → test → compare → promote or reject → log → repeat.

Round log0 entries
// 04PatternsFig. 4

Four loop architectures worth stealing.

The Champion Loop

Self-improving

The current version holds the title until a challenger beats it.

Sequence
Propose → test → compare → promote or reject → log → repeat.
Stop
Target · budget · 3 rounds without promotion.

Evidence — 7 rounds · $1.94 · holdout 3.4 → 4.6.

The Saturation Loop

Research

Batches of 20; every cluster keeps an original quote and source.

Evaluator
Does new evidence still teach anything?
Stop
Two empty batches · cap of ten.

Gate — a human traces every claim to a customer's words.

The Builder-Critic Loop

Adversarial

One model builds; another argues it is wrong. Objections are Fixed or Accepted on record.

Stop
No high-impact objection left unresolved.

Evidence — 11 objections → 7 changes · 4 accepted deliberately.

The Product Experience Loop

Experiential

Fresh browser session, real task, score the whole journey, fix the weakest screen, restart.

Guardrail
Pricing, legal, privacy, permissions require approval.

Evidence — "Invalid input." → named field, format, fix. Screen 2 → 4.

// 05Failure modesFig. 5

Five ways loops fail.

Shorter replies passed more checks. By round nine the answers were brief, polite, safe — and almost useless.

Fix

  • Restate the objective every round.
  • Multidimensional rubric.
  • Inspect behavior direction, not score.

If the judge is easier to fool than the task, the loop learns to please the judge.

Fix

  • Deterministic checks.
  • Evidence for judgments.
  • Test the evaluator on known cases.

One request becomes many rounds and passes. Rejected challengers still burned tokens.

Fix

  • Caps before execution.
  • Small models for simple checks.

A loop that works only while someone watches its terminal is still a prototype.

Fix

  • Async, resumable runs.
  • Notify on completion, failure, escalation.

Perfection unreachable → the loop spins and the budget bleeds.

Fix

  • Pair target with budget + stall rule.
  • Preserve best result; ask for the missing decision.
// 06EscalationFig. 6

Autonomy is earned, never assumed.

Stage 01 — Shadow

Runs on real work, affects nothing. Compare with the human decision.

BoundaryRead-only. Output: a shadow decision, logged next to the human one.

Q1Can the result be checked?

If you cannot explain how you would judge it yourself, clarify good work before automating it.

Q2What happens if it is wrong?

Automate what is checkable. Escalate what is consequential.

Q3Can you define the stop first?

"We will see how it goes" gives an unattended system no operating policy.

// 07ManifestoFig. 7

Six rules for reliable loops.

01

Spec before code

The blank shows which part of the system still needs design work.

02

One change per round

One change preserves causality.

03

The holdout is truth

Improvement-set wins are rumors until they survive the holdout.

04

Champion defends the title

A tie goes to the champion.

05

Checkable vs consequential

These need a human gate, not a better prompt.

06

Autonomy is earned

Shadow → suggest → approve → bounded, on observable evidence.

// 08The contractFig. 8 — one-page loop specification

The one-page loop specification.

Fill this in before building — the document on the right updates live on every click and keystroke.

Input · Fill every field — or admit the blank

Launch stage

Cartridge 07 · Loop Quest

CARTRIDGE 07 · LOOP QUEST (1986–2026) · SCROLL TO PLAY
  1. World 0-0 — MOST PEOPLE STILL USE AI ONE PROMPT AT A TIME.
  2. World 1-1 — YOU RAN THE LOOP IN YOUR HEAD: JUDGE, REMEMBER, ENFORCE, STOP.
  3. World 1-2 — WRITE THE SPEC. NINE PARTS. THE BLANK SHOWS WHAT NEEDS DESIGN.
  4. World 1-3 — CHAMPION DEFENDS ITS TITLE. ONE CHANGE PER ROUND. TIES TO CHAMPION.
  5. World 1-4 — FIVE MONSTERS: DRIFT, WEAK EVALS, COST, LATENCY, ENDLESS RETRIES.
  6. World 1-5 — AUTONOMY IS EARNED: SHADOW > SUGGEST > APPROVE > BOUNDED.
  7. World 1-6 — HOLDOUT 4.6 >= 4.5 - THE LOOP STOPS BY ITSELF. 40MIN. $2.
  8. World 1-7 — GAME COMPLETE. INSERT SPEC TO CONTINUE.