I stopped asking one model to write a whole piece about four months ago and my output roughly doubled. Not because I found a better model. Because I stopped asking any single model to do a job that has four completely different steps in it.

Writing is not one task. It is at least four, and they have almost nothing in common. Gathering raw material is a research job. Structuring it is an architecture job. Getting words on the page is a volume job. Making those words sound like a person is a taste job.

You would never hire one contractor to do all four. You would definitely never hire your most expensive contractor to do the volume step. And yet that is exactly what most people do every time they open a chat window and type "write me a 1,500 word article about."

Here is the split that works, what to run where, and the specific prompts that make the handoff clean.

The economics that force the split

Quick grounding, because this is the part that makes the whole thing obvious.

GPT-6 Astra shipped on September 3 at fifty dollars per million output tokens. Gemini 3.8 Flash sits roughly thirteen times cheaper per token. Fable 5.1 came out generally available on September 1 at the same fifty dollar output price as the version before it, though cache reads dropped to twenty five cents, which matters more than it sounds if you are running the same context repeatedly.

Now think about what a draft actually is. A draft is mostly volume. It is two thousand words of raw material, most of which you will cut, rearrange, or rewrite. You are paying premium output rates for text that exists to be destroyed.

The finish is different. The finish is maybe four hundred words of actual changed text, but every one of those changes requires judgment about rhythm, specificity, and whether a sentence sounds like a human wrote it. That is where the expensive model earns its keep.

Same piece. Two completely different jobs. One of them is worth thirteen times more per token than the other, and it is not the one you think.

Step one: raw material, and why you should not skip it

Nobody wants to do this step and it is the reason most AI writing is bad.

Before any model writes anything, you need raw material that only you have. Specific numbers. A thing a client said. What actually happened when you tried the thing. The objection you heard three times last month.

A model generating from nothing produces the average of everything it has read, which is the definition of generic. A model working from your material produces something only you could have published.

So step one is a dump, and you should do it out loud. Open a voice tool, talk for eight minutes about the thing, and let it transcribe. You will produce more usable raw material in eight minutes of talking than in forty minutes of typing, because talking bypasses the editor in your head that makes everything sound like a press release.

If the material lives in calls, pull it from there. I run Fathom across every call I take, which means I have a searchable archive of things real people actually said about real problems. When I need a specific objection in somebody's actual words, I go find it rather than inventing a plausible one. Invented objections read as invented. Every time.

The output of step one is a mess. That is correct. It is supposed to be a mess.

Step two: architecture, run at medium

This step is short and it is worth doing carefully because everything downstream inherits it.

You are asking a model to take your mess and propose a structure. Not to write. To organize.

The prompt that works:

"Here is raw material for a piece about [topic]. Do not write anything. Give me three possible structures for this piece. For each one, list the sections in order with a single sentence describing what each section does for the reader. Then tell me which of the three has the strongest opening and why. One of the three should be structured around a single argument. One should be structured around a process. One should be structured around a mistake and its correction."

Two things make this work. Forcing three options stops you from accepting the first plausible thing. And demanding three different shapes stops the model from giving you the same outline in three costumes, which is what it will do if you just ask for options.

Run this at medium effort. Structure is a reasoning task and it is short, so the cost is trivial and the quality difference is real.

Pick one. Or steal the opening from one and the middle from another, which is what I usually end up doing.

Step three: the volume draft, run cheap

Now hand the structure and the raw material to your cheapest capable model and tell it to write long.

The prompt:

"Write this piece following the structure exactly. Use the raw material provided. Where the material is thin, write the section anyway and mark it with [THIN] so I know to fill it. Write long rather than short. Do not polish. Do not try to be clever in the opening. Aim for 2,400 words. Use only the specific facts, numbers, and quotes in the material provided. Do not invent examples."

Three instructions in there are doing heavy lifting and I want to be explicit about them.

Write long rather than short. You want more material than you need because cutting is easy and generating is expensive. A 2,400 word draft you cut to 1,800 is better than a 1,800 word draft you have to expand.

Mark thin sections. This is the single most useful instruction I have added in a year. The model knows when it is padding. It just does not tell you unless you ask. Once you ask, it flags exactly the paragraphs where you need to go get more material, and those flags are almost always right.

Do not invent examples. Cheap models invent more than expensive ones. This instruction does not eliminate it but it cuts it substantially, and it makes the invention easier to spot because anything specific that you did not provide is now suspect by default.

Do not read this draft carefully. Skim it. You are checking that the structure held, not that the prose is good. It is not good. It is not supposed to be.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. Routing tables, gate logic, the scenario templates, and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

Step four: the finish, run expensive

Here is the part everybody gets wrong. They hand the draft to the good model and say "improve this."

That prompt produces a smoother version of the same thing. Smoother is not better. Smoother is usually worse, because the specific weird parts that made it sound human get sanded off in the name of flow.

You have to be specific about what the finish is for. Mine:

"Here is a draft. Do not restructure it and do not add new sections. Do these five things only.

One. Find every sentence that could have been written about any business and either make it specific to this one or delete it.

Two. Find the abstractions and replace them with the concrete thing underneath. If it says 'improved efficiency,' tell me what got faster and by how much, using only numbers already in the draft.

Three. Cut every sentence that restates the sentence before it.

Four. The opening paragraph should make a specific claim or tell a specific thing that happened. If it is throat clearing, replace it with the most interesting sentence currently buried in the piece.

Five. List every claim in the draft that is not supported by the source material, so I can verify or cut it.

Return the edited piece, then the list from item five separately."

Five constrained jobs beats one vague instruction. Every time. The list at the end is your hallucination check, and it is the reason I sleep fine publishing work that started with a cheap model.

Step five: the pass only you can do

The model cannot do this one and you should stop hoping it will.

Read the whole thing out loud. Actually out loud, not in your head. Every place you stumble is a place a reader stumbles. Fix those.

Then find three places to put something the model could not have written. A thing you actually think. A joke that risks something. An admission that you got it wrong for two years. That is the whole difference between content that gets read and content that gets scrolled past, and it takes about six minutes.

Then cut the last paragraph. It is almost always a summary of what you just said, which nobody needs, because they just read it. End on the sharpest thing you have instead.

Where this saves the actual time

People assume the savings is money. The money is real but it is not the point.

The savings is that steps three and four are the only ones that take a model any real time, and they now happen while you are doing something else. You do the eight minute talk, spend five minutes picking a structure, kick off the volume draft, and go take a call. Come back to a draft. Kick off the finish. Go take another call. Come back to something publishable.

The human work compresses to about twenty five minutes of real attention spread across a morning, instead of ninety minutes of sitting in a chair fighting a blank page.

That is the whole game. Not fewer minutes of work. Fewer minutes of your attention held hostage.

Wiring it so you do not have to think about it

Once the four prompts are stable, stop copying and pasting them.

The version I run chains through Make.com. A transcript lands in a Drive folder. The scenario picks it up, runs the structure prompt, writes three options to a doc, and waits. I pick one by putting an X in a cell. It runs the volume draft against my cheap model, then the finish prompt against my expensive one, and drops both versions plus the unsupported claims list into a doc named for the piece.

Total human input, three decisions. Structure choice, unsupported claims review, and the read aloud pass.

If you write across several brands and voices, run separate finish prompts per brand rather than one prompt with a voice section. Voice instructions get diluted when they compete with five other rules. Separate prompts stay sharp.

For managing where the finished pieces go, Buffer handles the distribution end well enough that I stopped thinking about it, and if the destination is a newsletter, Beehiiv will take the HTML directly without a fight.

The objection I get every time

"Doesn't running two models make it worse than just using the best one?"

No, and the reason is that the best model is not best at everything. It is best at judgment. It is not meaningfully better at producing two thousand words of scaffolding from a structure you already specified, because that task does not require judgment, it requires typing.

You are not degrading the output. You are removing the expensive model from a job where its advantage does not apply, and concentrating it entirely on the job where its advantage is the whole point.

The test is easy if you do not believe me. Run one piece both ways. Full pipeline versus one model doing everything. Read them side by side tomorrow morning, when you have forgotten which is which. Then go with whichever wins.

I have run that test maybe fifteen times. The pipeline has won every time except twice, and both losses were pieces where I skipped step one and gave it nothing to work with. Which tells you where the real leverage is, and it is not in the models at all.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

Do this with your next piece

Do not rebuild your whole process. Take the next thing you have to write and change exactly one thing.

Before you open any model, talk for eight minutes about the topic and transcribe it. Then feed that transcript in along with whatever you would normally have typed as a prompt.

That single change, by itself, does more for the output than any model upgrade you will make this year. Everything else in this piece is optimization on top of it.

Jordan

The AI Newsroom | Practical AI for people with a business to run.