Bonus world · Loop engineering

The Run Sheet

“Not about how models work inside, but about what you build around them. We follow one real agent run around a desk, from the failing test that starts it to the check that catches a bad fix.”

Unpacking the world: 1.2 MB of drawings and measured data.

AgenticAmit A field guide to loop engineering

Field guide · loop engineering

Stop typing prompts. Build the loop.

An AI agent is a model called again and again. Loop engineering is everything around those calls: what starts a run, what the model sees, what it's allowed to do, how the work is checked, and when it stops.

Scroll to follow one real run ↓

Fig. 01 · the definitionJun 7, 2026
Addy Osmani: Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.

Addy Osmani · Loop Engineering

So who's typing the prompts?

Pinned · what a loop isJun 30, 2026
Claude blog: we define loops as agents repeating cycles of work until a stop condition is met.

Claude · Getting started with loops

Pinned · the model's jobDec 19, 2024
Anthropic: They are typically just LLMs using tools based on environmental feedback in a loop.

Anthropic · Building effective agents

Pinned · the stop ruleAug 2026 · preprint
arXiv 2608.21884: These systems start agent runs on a schedule or on repository events and stop them when a machine-checkable condition holds.

Lulla et al. · arXiv 2608.21884

Fig. 03 · the smallest loop there isJul 14, 2025
Geoffrey Huntley: Ralph is a Bash loop. while :; do cat PROMPT.md | claude-code ; done

Geoffrey Huntley · Ralph

01 · The triggerCI

✕ test_bill.py
2 failed, 2 passed

Run started →
02 · The pagefor line 01

        
03 · The modelrented, not built
reads the page ·
writes one step
04 · The tool callthe menu
  1. run_tests()
  2. read_file(path)
  3. edit_file(path, text)
05 · The worldsandbox · a walled-off copy

        
06 · The checka machine decides

tests all pass? → exit
lines left? → go again
none left? → a person

Exit · the answerwith its evidence
outbox
Hand-off · a personbelow the budget rule
If the budget runs out, the sheet lands here.
Fig. 02 · run sheetan illustrative run
test_bill.pymake it pass

started by · CI failure

    Budget · 6 linespast this rule, a person takes over

    Lines lefttests 2/4·
    The answer · 4 passed + cents = round(total * 100)
    + base, extra = divmod(cents, people)
    + shares[-1] += extra
    You designwhat starts a run You designwhat goes on the page You designwhat it may do You designwho decides it's done You designthe budget You designwho takes over

    Before line 01

    Nobody opened a chat window.

    A test failed in CI, the robot that runs your tests every time code changes. That failure started this run on its own. The job: fix a little app that splits a dinner bill between friends.

    Follow the run around the desk.

    02 · the page

    Everything it knows fits on one page.

    The model remembers nothing between calls. So each turn, the code around it (called the harness) writes it a fresh page: rules, task, tools, and everything that's happened so far. That page is the context.

    TakeawayIf it isn't on the page, it doesn't exist.

    03 · the model

    It reads the page and writes one next step.

    The model is the one part you rent instead of build. It can't run anything. It reads, and writes text back. This time the text is a plan: see it fail first.

    PinnedAnthropic: agents are "typically just LLMs using tools based on environmental feedback in a loop."

    04 · the tool call

    A tool call is a request, not an action.

    It picked from a menu you wrote: three tools, each with a name and a shape. It asked to run the tests. Nothing has happened yet. Your code decides whether, and how, to do it.

    TakeawayA few clear tools beat many vague ones.

    05 · the world

    Your code does the doing.

    The harness runs the tests inside a sandbox, a walled-off copy of the project, and captures what comes back, exactly as it came back. It's the only moment in the turn that touches anything real.

    TakeawayThe result is what the model reads next. Keep it short and true.

    06 · the check

    Done, or go again?

    A machine makes this call, not the model. All tests green: stop and hand back the answer. Otherwise go again, unless the budget is spent. Then the loop stops and a person takes over.

    PinnedThe research review: runs stop "when a machine-checkable condition holds."

    The run sheet

    Everything you just watched is one line.

    Trigger, page, model, tool call, world, check: that was one turn. The harness writes each turn onto the run sheet as one line. Watch the end of the line.

    The carriage return

    What happened goes on the next line.

    Think of an old typewriter. At the end of a line, the carriage swings back to the left margin, one line down. Here, the test result rides that return onto line 02's page. That's the loop.

    ▸ the same idea, as LangChain puts it

    LangChain: The verification loop adds a grader: something that checks the agent's output against a rubric and, if it fails, sends the result back with feedback.

    LangChain · The Art of Loop Engineering · Jun 16, 2026

    Line 02 · it types by itself

    Now nobody's pressing keys.

    Line 02's page includes the failure, so the model asks to read the code. It finds the maths: each share is rounded to the cent, so three shares of $33.33 add up to $99.99. A cent goes missing.

    Line 03 · a fix

    A fix that looks right.

    Give the last person whatever's left over. On paper, that's perfect. And because the rules say so, the harness reruns the tests after every edit.

    Line 03 · the catch

    The check caught it. Not the model.

    Computers store numbers in binary, and some decimals can't be written exactly. So 70 minus two shares of 23.33 comes out as 23.340000000000003. The test fails. That's the sandbox, exactly as observed.

    Line 03 · the return

    The failure rides back down.

    That exact line, 23.340000000000003 != 23.34, is carried to the left margin and lands on line 04's page. The next fix will be written with it in view.

    Line 04 · the fix

    Count in whole cents.

    So the next fix is different: do the maths in whole cents, where nothing can round wrong, and give the leftover cent to the last person.

    +    cents = round(total * 100)+    base, extra = divmod(cents, people)+    shares[-1] += extra$ pytest -q   4 passed in 0.03s

    Exit

    It leaves with its evidence.

    Four tests pass, so the check says done. The answer goes out through the right margin with the change and the test run attached, so a person can see why it's trusted.

    After line 04

    Four lines. Two to spare.

    One bad fix, caught by the loop itself, and two lines of budget left over. The sheet is also a record: every page, call and result is still there for a person, or the next loop, to read.

    Illustrative runEvery tool output is real (pytest 9.1.1, Python 3.14, 27 Sep 2026). The model's choices were written for this page.

    What you actually build

    The model wrote four lines. You built the desk.

    What starts a run?A failing check, a schedule, a new issue.Would I be fine with this starting at 3am?
    What goes on the page?Rules, task, tools, history, trimmed.What does it need that it can't see?
    What may it do?A short tool menu, run in a sandbox.What's the worst this tool could do?
    Who decides it's done?Tests, a linter, a grader. Not the model.Could this pass while still being wrong?
    When does it give up?A cap on lines, tokens or money.What is one run worth to me?
    Who takes over?A person, with the whole sheet attached.Who gets pinged, and what do they see?

    The loop above

    And there's a loop above this one.

    Finished sheets are a record of what worked and what didn't. LangChain describes reading them to improve the rules, the tools and the check for the next run, and calls it a hill-climbing loop. The outbox feeds the page.

    Where it started

    One line of Bash is already a loop.

    Geoffrey Huntley's "Ralph" runs the agent on the same prompt file, forever. Everything else on this desk is what you add so it knows when it's done, when to stop, and who to hand it to.

    Take this with you

    The model writes one line at a time. You engineer the sheet.

    Pick one task you keep re-prompting by hand. Write down what would start it, what goes on the page, how a machine would know it's done, and when it should stop and ask you. That's your first loop.

    Amit, thumbs up