Nobody opened a chat window.
A test failed in CI, the robot that runs your tests every time code changes. That failure started this run on its own. The job: fix a little app that splits a dinner bill between friends.
Follow the run around the desk.

Bonus world · Loop engineering
The Run Sheet
“Not about how models work inside, but about what you build around them. We follow one real agent run around a desk, from the failing test that starts it to the check that catches a bad fix.”
Unpacking the world: 1.2 MB of drawings and measured data.
An AI agent is a model called again and again. Loop engineering is everything around those calls: what starts a run, what the model sees, what it's allowed to do, how the work is checked, and when it stops.
Scroll to follow one real run ↓
Addy Osmani · Loop Engineering
So who's typing the prompts?
Claude · Getting started with loops
Anthropic · Building effective agents
Lulla et al. · arXiv 2608.21884
Geoffrey Huntley · Ralph
✕ test_bill.py
2 failed, 2 passed
tests all pass? → exit
lines left? → go again
none left? → a person
started by · CI failure
Budget · 6 linespast this rule, a person takes over
A test failed in CI, the robot that runs your tests every time code changes. That failure started this run on its own. The job: fix a little app that splits a dinner bill between friends.
Follow the run around the desk.

The model remembers nothing between calls. So each turn, the code around it (called the harness) writes it a fresh page: rules, task, tools, and everything that's happened so far. That page is the context.
TakeawayIf it isn't on the page, it doesn't exist.
The model is the one part you rent instead of build. It can't run anything. It reads, and writes text back. This time the text is a plan: see it fail first.
PinnedAnthropic: agents are "typically just LLMs using tools based on environmental feedback in a loop."
It picked from a menu you wrote: three tools, each with a name and a shape. It asked to run the tests. Nothing has happened yet. Your code decides whether, and how, to do it.
TakeawayA few clear tools beat many vague ones.
The harness runs the tests inside a sandbox, a walled-off copy of the project, and captures what comes back, exactly as it came back. It's the only moment in the turn that touches anything real.
TakeawayThe result is what the model reads next. Keep it short and true.
A machine makes this call, not the model. All tests green: stop and hand back the answer. Otherwise go again, unless the budget is spent. Then the loop stops and a person takes over.
PinnedThe research review: runs stop "when a machine-checkable condition holds."
Trigger, page, model, tool call, world, check: that was one turn. The harness writes each turn onto the run sheet as one line. Watch the end of the line.
Think of an old typewriter. At the end of a line, the carriage swings back to the left margin, one line down. Here, the test result rides that return onto line 02's page. That's the loop.
▸ the same idea, as LangChain puts it
LangChain · The Art of Loop Engineering · Jun 16, 2026
Line 02's page includes the failure, so the model asks to read the code. It finds the maths: each share is rounded to the cent, so three shares of $33.33 add up to $99.99. A cent goes missing.
Give the last person whatever's left over. On paper, that's perfect. And because the rules say so, the harness reruns the tests after every edit.
Computers store numbers in binary, and some decimals can't be written exactly. So 70 minus two shares of 23.33 comes out as 23.340000000000003. The test fails. That's the sandbox, exactly as observed.

That exact line, 23.340000000000003 != 23.34, is carried to the left margin and lands on line 04's page. The next fix will be written with it in view.
So the next fix is different: do the maths in whole cents, where nothing can round wrong, and give the leftover cent to the last person.
+ cents = round(total * 100)+ base, extra = divmod(cents, people)+ shares[-1] += extra$ pytest -q 4 passed in 0.03sFour tests pass, so the check says done. The answer goes out through the right margin with the change and the test run attached, so a person can see why it's trusted.
One bad fix, caught by the loop itself, and two lines of budget left over. The sheet is also a record: every page, call and result is still there for a person, or the next loop, to read.
Illustrative runEvery tool output is real (pytest 9.1.1, Python 3.14, 27 Sep 2026). The model's choices were written for this page.
Would I be fine with this starting at 3am?
What does it need that it can't see?
What's the worst this tool could do?
Could this pass while still being wrong?
What is one run worth to me?
Who gets pinged, and what do they see?
Finished sheets are a record of what worked and what didn't. LangChain describes reading them to improve the rules, the tools and the check for the next run, and calls it a hill-climbing loop. The outbox feeds the page.
Geoffrey Huntley's "Ralph" runs the agent on the same prompt file, forever. Everything else on this desk is what you add so it knows when it's done, when to stop, and who to hand it to.
Pick one task you keep re-prompting by hand. Write down what would start it, what goes on the page, how a machine would know it's done, and when it should stop and ask you. That's your first loop.