Max Woolf published something last week that's worth more to you than most model launches.

He pointed AI agents at handwritten Rust code and told them to make it faster. Not "improve it." Not "clean it up." Make it faster, measured by a benchmark, and prove it still works, measured by a test suite. Then run again. And again.

The loops produced speedups of roughly 2 to 20 times on individual workloads, and 7.5 to 32 times cumulatively across generations.

That's not the interesting part, though. The interesting part is his finding about when the loops worked and when they didn't. They worked when there was a number to push down and a test that proved the result was still correct. They wandered when there wasn't.

You're not writing Rust. I'm not either. But that finding applies to almost every piece of repeat work you hand to a model, and most people are doing it wrong in the exact way Woolf's result predicts.

The loop most people are running

Here's the normal pattern. You ask for something. It comes back. You say "make it better." It comes back different. You say "better, but shorter." It comes back shorter and a little worse in some other way. You say "no, keep the part about pricing." Six rounds later you've got something acceptable and a vague feeling that round three was actually the best one.

That's a loop with no number and no test. The model is guessing what "better" means every single time, and you're grading on feel. It wanders because there's nothing to hold it in place.

Now look at what Woolf did. Speed is a number. The test suite is a pass or fail. Each round, the model either made the number go down while keeping the tests green, or it didn't. No guessing. No feel. The loop can run on its own because the judge isn't a person's mood.

You can build the same structure for business tasks. Not every task, but more than you'd expect.

The four parts of a loop that works

Every good loop needs four things, and you write all of them before you start.

The target. One number, and which direction you want it to go. Word count down. Reading grade level down. Number of required points covered up. Cost per run down. Time to complete down. One number. If you have two, pick the one that matters and turn the other into a test.

The test. A list of things that must stay true no matter what. These are pass or fail, not scored. "Mentions the price." "Includes the refund policy." "Keeps the client's name spelled correctly." "Returns valid JSON with all six fields." If any test fails, that round is thrown out, no matter how good the number looks.

The stop. When the loop ends. A maximum number of rounds, a target number that counts as done, or "stop when two rounds in a row fail to improve." Without a stop, loops either end too early because you get bored, or run forever producing smaller and smaller changes.

The log. A record of every round: what changed, what the number was, whether the tests passed. This is what lets you pick the best round instead of the last round, which aren't always the same thing.

The loop prompt

Here's the base prompt. It works in any chat interface for manual loops, and it's the same logic you'd put into an automation.

❝

You are running an improvement loop on the text below. Your job is to change it to improve one measurable target while keeping every requirement true.

TARGET: [the number, and which direction]

REQUIREMENTS: every one of these must remain true after your change. If a change would break any of them, don't make it.

1. [requirement]

2. [requirement]

3. [requirement]

For this round:

First, state the current value of the target.

Second, make one focused change. Not a full rewrite. One change that moves the target.

Third, state the new value of the target.

Fourth, check every requirement and state pass or fail for each one, quoting the text that satisfies it.

Fifth, output the full revised text.

If you can't find a change that improves the target without breaking a requirement, say so and stop.

That last line matters more than it looks. It gives the model permission to say "this is as good as it gets," which stops it from making pointless changes just to have something to show you.

Five loops worth running this month

Here are five places I run this in my own business. Each one has a real number and a real test.

Loop one: shrinking a system prompt. This one pays for itself right now. If you've got a long system prompt running in an automation, every token of it gets billed on every single call. Target: token count, down. Test: your saved set of real test inputs, which all still have to produce acceptable outputs. I cut one of mine from about 2,400 tokens to 1,100 in six rounds, with every test still passing. On a workflow that runs a few thousand times a month, that's real money, even at last week's lower prices.

Loop two: tightening a sales email. Target: word count, down. Tests: must mention the specific outcome, must include the price or the next step, must name the prospect's industry, must end with a single clear ask. Most sales emails are forty percent longer than they need to be. This loop finds the forty percent.

Loop three: simplifying client instructions. Onboarding documents, how to guides, setup steps. Target: reading grade level, down. Tests: every step still present, in order, every link still included, every warning still included. Your clients will actually read the result, which fixes more support tickets than any FAQ page.

Loop four: covering the objections. This one runs in the other direction. Target: number of common objections addressed, up. Test: total length stays under a limit you set. Give it your list of the eight objections you hear most, and your current sales page copy. Each round it works one more objection in without blowing out the length.

Loop five: speeding up a spreadsheet. If you've got a slow Google Sheet full of nested formulas, this is Woolf's experiment almost exactly. Target: calculation time or number of volatile functions, down. Test: output values match the original on a fixed set of rows. Copy the sheet, run the loop on the copy, compare the outputs cell by cell.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

Automating the loop

Running a loop by hand works fine for five or six rounds. For anything longer, or anything you'll repeat, automate it.

In Make.com the structure is a repeater with a counter. Each pass sends the current text to the model with the loop prompt, parses the response, checks whether the tests passed and whether the target improved, and writes a row to a log sheet. If both are true, the new text becomes the current text for the next round. If not, it keeps the previous version and tries again. A router at the top checks the stop condition.

Here's the important detail most people miss. Don't let the same model that makes the changes be the only judge of whether the tests passed. Models are generous graders of their own work. Where you can, check the tests with plain logic instead. Word count, token count and "does this text contain the word refund" are all things a formula can check with zero opinion involved. Save the model judgment for the tests that genuinely need it.

And for the tests that do need a model, use a different one than the one doing the edits. Having two or three models one login apart, through something like Galaxy.ai, makes this easy. One model edits, a different one checks, and the plain logic checks whatever it can.

Where loops go wrong

Three failure modes to watch for.

Gaming the number. Loops are very good at finding the cheapest way to move a number, and the cheapest way is often the wrong way. Tell a loop to reduce word count and it will happily turn your warm opening into a bullet list. That's why the tests exist. If you see the number improving and the output getting worse, you're missing a test. Add it.

Death by tiny changes. After a few rounds, the gains get small. Round eight saves four words. Round nine saves two. That's your stop condition telling you it's done. Set "stop after two rounds with less than a two percent improvement" and you'll save yourself a lot of pointless runs.

Losing the voice. Every round of editing pulls the text a little closer to how the model writes by default. After ten rounds, your email can sound like everyone else's. The fix is a voice test: "keeps at least three sentences under eight words" or "uses first person throughout" or "includes one line of dry humor." It feels silly to write that as a test. It works.

Where loops don't belong

This isn't for everything, and knowing when not to use it saves you a lot of wasted runs.

Don't loop on anything where "better" is genuinely a matter of taste with no measurable piece. Brand voice, a big creative idea, a new offer concept. Those need a person with judgment and a few honest conversations, not a counter and a log.

Don't loop on anything where the test is "is this true." Loops can't verify facts. They can check that a claim is present, not that it's correct.

And don't loop on something you'll only use once. The setup cost, writing the target and the tests, is worth it for things you'll run again. For a one time email, just write it well.

The deeper habit

The real value here isn't the loop. It's what writing the loop forces you to do.

Most of the time when you're frustrated with AI output, the problem isn't the model. It's that you haven't decided what "good" means in a way anyone could check. You know it when you see it. That's not enough for a model, and honestly it's not enough for a new employee either.

Writing a target and a set of tests turns "I know it when I see it" into something concrete. Once you've done that for a task, you can hand that task to a model, a contractor or a team member and get consistent results. The loop is just the most mechanical way of cashing in on that clarity.

Woolf's finding comes down to this: agents are very good at pushing a number when you give them one. The work on your end is deciding which number matters.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

This week

Find the longest system prompt you have running in any automation.

Write down six test inputs from real work, and what an acceptable output looks like for each.

Run the loop prompt with token count as the target and those six tests as the requirements. Stop after five rounds.

Then check the cost difference per month. That number is usually big enough to make you go looking for your second longest prompt the same afternoon.

Jordan

The AI Newsroom | Practical AI for people with a business to run.