I watched somebody spend forty minutes on an image prompt last week.
Not forty minutes generating. Forty minutes writing. Adding clauses, moving adjectives around, specifying that the laptop should be on the left side of the desk and angled slightly toward the viewer and not too close to the coffee cup, which should be behind it and to the right but not so far right that it leaves the frame.
Four hundred words describing a picture. Which, if you think about it for one second, is an insane thing to do, because a five year old could have communicated the same information in eleven seconds with a crayon.
On September 8, ChatGPT Images 2.5 shipped with a Sketch feature. You hand it a drawing, it hands you back an image. Latency dropped by about half on top of that, and the API split into two flavors: Flare, which is fast and good enough, and Sunburst, which is slower and more detailed.
The sketch part is the part that changes how you work. And almost nobody's using it, because everybody's still in the habit of typing.
Why your prompts keep missing
Here's the thing nobody says out loud about image prompting.
Language is sequential. You say one thing, then the next thing, then the next. Spatial relationships are not sequential. They're simultaneous. Everything in a composition exists at the same time in relationship to everything else.
When you describe a layout in words, you're converting a two dimensional thing into a one dimensional stream and hoping the model reconstructs it. That's a lossy conversion and it gets lossier the more elements you add. Two objects, words work fine. Five objects with specific relative positions, you're now writing a legal document and the model is still going to put the coffee cup in the wrong place.
The failure isn't the model's. It's the encoding. You picked a bad format for the information.
A sketch is the right format. It carries position, scale, relationship and emphasis all at once, natively, with zero ambiguity, and it takes ninety seconds to make.
The ninety second sketch
You're not drawing. Let's be clear about that, because the word "sketch" makes people who can't draw close the tab.
You're making a diagram. Boxes and labels. That's the entire skill.
Here's the method:
Draw the frame first. A rectangle in the correct aspect ratio. If you're making a 1200 by 630 banner, draw a wide rectangle. If it's a 1080 by 1080 square, draw a square. This one step fixes more composition problems than everything else combined, because the model now knows what shape it's filling.
Block in the big shapes. Rough rectangles and circles where things go. A rectangle for the laptop. A circle for the cup. A blob for the person. Do not add detail. Detail is the model's job and adding it in a bad drawing actively hurts, because now the model is trying to reproduce your terrible perspective.
Label everything. Write the word inside or next to the shape. "laptop." "coffee, steaming." "window, city at night." The labels carry the semantics, the shapes carry the geometry, and between them you've said everything.
Mark the focal point. A star, a circle, an arrow, whatever. One mark on the thing the eye should land on first. This is the single highest value mark on the page and I've never seen anybody make it unprompted.
Arrow the direction. If somebody's looking somewhere, or something's moving, draw the arrow. Direction is nearly impossible to nail in prose and trivial to draw.
Phone camera, napkin, iPad, whiteboard, doesn't matter. I run image work through Galaxy.ai mostly so I can throw the same sketch at two or three image models without juggling subscriptions, and the differences between them on layout adherence are bigger than you'd expect. Photograph it badly. Upload it. The model is better at reading bad drawings than you'd believe, and it's a lot better at reading a bad drawing than a good paragraph.
The three shot workflow
One sketch and one prompt still won't get you there. Here's the sequence that does.
Shot one: structure. Upload the sketch. Prompt for layout only. Something like: "Follow this layout exactly. Placement, scale and relationships should match the sketch. Render as a clean photographic scene. Ignore my line quality, the drawing is a diagram not a reference." You're judging one thing in the output: is everything where it belongs? Nothing else matters yet.
Shot two: style. Once the structure is right, keep the image and change the treatment. "Same composition. Shift to a darker palette, cooler light coming from the left, more contrast." You're only moving the look. Composition is locked, so you can be aggressive here without breaking what already works.
Shot three: fix. One specific correction. "The cup should be smaller and further back." One change per shot. The instinct to fix four things at once is the instinct that costs you nine more generations.
Three shots, maybe five if you're picky. Versus the forty minute prompt that never quite lands. I've watched this cut image production time by roughly two thirds on repeatable formats, and the bigger win isn't speed, it's that the output actually matches what was in your head.
FROM THE AI NEWSROOM
The AI Workflow Blueprint
The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.
Get the Blueprint for $47
.....
Where this actually pays
Not every image is worth this. Here's where the sketch approach earns real money.
Ad creative variants. This is the big one. You have a winning ad. You want six structural variants, not six color swaps. Draw six layouts in ten minutes, run each through the three shot workflow, and you've got a test set that varies the thing that actually moves performance, which is composition and hierarchy, not whether the button is teal.
Worth noting that the platforms are moving toward this too. OpenAI's Ads Manager now suggests copy and imagery off your landing page and campaign objective, and Meta's been pushing automated creative generation hard. Those tools produce competent, average output. The reason to keep a hand on the wheel is that average is what everybody else's automated tool produces too.
Diagrams and explainers. Anything where you need to show a process, a stack, a before and after. Words are terrible at this and stock illustration is worse. A labeled sketch turned into a clean rendered diagram is the single most underused content asset in small business marketing.
Thumbnails. Face position, text zone, focal object, contrast block. All spatial. All impossible to specify in prose and trivial to sketch. If you're making video thumbnails and you're not sketching first, you're leaving click through on the table.
Product placement mockups. Show the thing in the context it'll be used. Draw the room, draw the box where the product goes, label everything else.
Carousel layouts. Ten slides with consistent structure. Sketch the master once, reuse it, change only the labels. Then load the set into Buffer and the whole week ships in one sitting.
Video thumbnails and openers. Same sketch, and if you're producing talking head video from scripts, HeyGen handles the footage while the sketch handles the frame around it.
Flare or Sunburst
Quick and practical, because the temptation is to always reach for the better one.
Use Flare for everything in the structure phase. You're checking placement, not admiring the render. Fast and cheap is exactly right, and when you're doing three shots per asset across eight assets, the difference adds up to real money and real minutes.
Use Sunburst on the final pass only, and only when the image is going somewhere that deserves it. A paid ad, a landing page hero, a printed piece. For a blog header that'll be seen at 600 pixels wide on a phone, Flare's output is indistinguishable and costs a fraction.
The rule I use: if I'm going to look at it more than once, Sunburst. If it's going to scroll past, Flare.
Locking your brand look
The problem with generated imagery for a business is consistency. One image looks great. Forty images look like forty different companies.
Two things fix this and both are cheap.
The reference sheet. One image, made carefully, that represents your visual language. Palette, lighting direction, texture, level of realism. You attach it to every generation alongside the sketch and say "match the visual treatment of the attached reference, follow the layout of the attached sketch." Two images in, one image out, consistency solved.
The locked sketch template. For each format you produce regularly, make the sketch once and save it. Your 1200 by 630 newsletter banner has a masthead zone, a headline zone, a focal zone and a footer bar. Draw that once. From then on you're not sketching, you're relabeling, which takes about twenty seconds.
I keep both in a Notion page per brand with the reference image at the top and the sketch templates underneath, and I've stopped having the "does this look like us" conversation entirely.
What it still can't do
Be realistic about where the wall is, because hitting it unexpectedly wastes an hour.
Typography is still not reliable. It'll produce text that looks like text and often reads correctly at a glance, but it will not match your brand typeface and it will not hold up to somebody actually reading it. Generate the image without text, add text in a real design tool. Every time. No exceptions.
Exact brand color is still a coin flip. It'll get close. Close isn't your hex. If brand color matters, produce the image in a neutral palette and color grade after.
Dense fine detail at small scale degrades. Logos on products, text on signage in the background, patterns with specific repeats. Plan around it rather than fighting it.
And it can't make a decision for you. It'll happily generate forty variations of an idea that was weak to begin with. The sketch step is partly valuable because it forces you to decide what the image is doing before you spend anything on generating it, and about a third of the time the honest answer at that stage is "this image doesn't need to exist."
The habit underneath all of this
Here's the broader point and then I'll get out of the way.
The instinct with these tools is always to add more language. More detail, more qualifiers, more specificity. And there are places where that works beautifully. Text output, analysis, code, the more precisely you specify the better you do.
Images aren't that. Images are a place where the right move is to stop typing and switch formats. Ninety seconds with a pen beats four hundred words every single time, and the reason people don't do it isn't that it's hard. It's that it feels unserious. It feels like you're not really working.
You are. You're just using the right encoding for the job, which is most of what good work is.
FROM THE AI NEWSROOM
The AI Business Accelerator
For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.
Join the Accelerator for $97
.....
This week
Find the last image you generated that didn't come out right. You know the one.
Draw it. Frame, blocks, labels, focal mark, arrows. Give yourself ninety seconds and don't go over.
Photograph it, upload it, one prompt about structure only. Compare that to what you got from the paragraph.
If it's better, and it will be, go make your locked sketch template for whatever format you produce most. That's the twenty minutes that pays for the next six months.
Jordan
The AI Newsroom | Practical AI for people with a business to run.

