Two pieces of research came out last week that, read together, explain most of the frustration people have with long AI work sessions.
The first came from Meta's FAIR lab. In a paper on what they call context language models, they found that a single instruction about how to manage context improved performance on a held out task by as much as 35.9 points. That's one instruction. No new model, no extra compute. Just a sentence telling it what to hang on to and what to toss.
The second was a dataset called TraceML. Researchers collected thousands of real human work sessions from Kaggle competitions and compared them to AI agent runs on the same kind of problems. The headline finding: agents collapse into narrow loops. They keep trying small variations of the same approach. Humans pivot to a genuinely different approach about three times as often.
If you've ever spent an hour in a chat with a model on a hard problem, you've lived both of these.
The conversation starts great. Somewhere around the twentieth message, things get mushy. The model forgets a constraint you set at the beginning. It starts contradicting a decision you made together. It keeps proposing tweaks to an idea that isn't working instead of stepping back. You end up re explaining things you already explained, and eventually you start a new chat and lose everything you'd built up.
That's a context problem and a pivot problem. And both have prompt level fixes.
This week's vault is about cleaning the desk. The prompts that keep long work sessions sharp from the first message to the fiftieth.
Why long sessions rot
Think of the model's context as a desk.
At the start of a session it's clean. Your instructions are right there in front of it. As you work, everything piles up: drafts, rejected ideas, side questions, corrections, the thing you pasted in to check one detail, three versions of the same paragraph.
The model can technically see all of it. But the important stuff, the original goal, the constraints, the decisions you made, gets buried under the working mess. The model starts weighting recent clutter over early instructions. Forgetting is the wrong word for it. The signal is drowning in noise.
Humans do the same thing, by the way. That's why good operators keep a notebook with the decisions written down instead of trusting their memory of a three hour meeting.
The fix is to make the model keep that notebook.
Prompt one: the working state
This goes at the very start of any session you expect to run long. A proposal you're building, a strategy document, a complex analysis, a project plan.
Before we start, set up a WORKING STATE block that you'll maintain for this whole conversation. It has four sections:
GOAL: what we're trying to produce, in one or two sentences.
CONSTRAINTS: every rule, requirement or limit I've given you.
DECISIONS: every choice we've made together, with a few words on why.
OPEN QUESTIONS: anything unresolved.
Every time something changes, a new constraint, a decision, a question answered, update the block and show it at the end of your response. Keep it short. If the block and something earlier in the conversation disagree, the block wins.
Don't skip that last line. When the desk gets messy, it tells the model which version to believe. Instead of searching back through forty messages to figure out whether you decided on the $2,500 or the $3,000 package, it looks at the block.
The other benefit is for you. At any moment you can glance at the block and see whether the model actually understood. If a decision is missing or wrong, you catch it immediately instead of three drafts later.
Prompt two: the reset
Even with a working state, long sessions accumulate junk. Every so often, clear the desk.
Let's reset. Write a handoff memo that someone with no access to this conversation could use to continue the work. Include: the goal, every constraint, every decision and why, the current state of the work, what's been tried and rejected and why, and the next step. Then stop.
Copy the memo. Open a fresh chat. Paste it in as the first message.
You've just thrown away all the clutter and kept everything that matters. The new session starts with a clean desk and a perfect briefing. In practice the quality jump is obvious, often within the first response.
I do this about every twenty to thirty messages on anything serious. Sure, it feels like extra work. Then you remember the twenty minutes you'd otherwise spend wrestling a session that's gone soft.
There's a bonus here, too. Those handoff memos are the best project documentation you'll ever have. Save them. If you work in Claude or any tool that supports projects with stored documents, drop each memo into the project files and every future chat on that work starts from it.
Prompt three: the pivot rule
Now the TraceML problem. Models get stuck making small variations on an approach that isn't working. Ask for a better headline and you'll get the same headline with a different adjective, nine times in a row.
You fix this by giving the model an explicit rule for when to change direction.
Pivot rule for this session: if we've tried the same general approach twice and I'm still not happy, don't make a third variation. Instead, stop and propose three genuinely different approaches, each starting from a different assumption. For each one, say in one sentence what it assumes that the current approach doesn't. Then let me pick.
The key phrase is "starting from a different assumption." Without it, you'll get three variations dressed up as three approaches. With it, you get actual alternatives.
Here's what that looks like in practice. Say you're stuck on a sales email that isn't landing. The current approach assumes the prospect needs to understand your process. The pivot might offer one version that assumes they already trust you and just need a reason to act now, one that assumes the real buyer is their boss and writes for forwarding, and one that assumes they're not ready and offers something small instead of a call.
Those are three different emails, not three edits. That's what a human strategist does when they're stuck. They question the premise.
Prompt four: the drift check
Use this one whenever something feels off. The model's suggestions are drifting, or it just did something that contradicts what you agreed.
Pause. Compare your last two responses to the GOAL and CONSTRAINTS in the working state. List anything that drifts from the goal or breaks a constraint. Then tell me whether we should fix the drift or update the working state because the goal has genuinely changed.
That last choice matters. Sometimes the drift is the model getting lost. Sometimes the drift is the work telling you the goal was wrong. Turning on purpose is fine. Wandering off isn't. This prompt makes you call which one you're doing.
FROM THE AI NEWSROOM
The AI Workflow Blueprint
The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.
Get the Blueprint for $47
.....
Prompt five: the attic
Long sessions produce a lot of good material that doesn't fit the current goal. An idea for a different offer. A phrase that's great but wrong for this piece. A side analysis.
If you leave it in the conversation, it's clutter. If you delete it, it's gone. Give it somewhere to live.
Keep an ATTIC section below the working state. Whenever we set aside something worth keeping that doesn't fit the current goal, add one line to the attic describing it. Don't bring attic items back into the work unless I ask.
At the end of the session, the attic is a list of everything worth revisiting. Copy it somewhere you'll actually see it. Some of my best ideas this year started as attic lines from sessions about something else entirely.
Putting it together
Here's how these work as a system on a real piece of work. Say you're spending the afternoon building next quarter's offer and pricing.
Message one: the working state prompt, plus the pivot rule and the attic instruction, plus your actual brief. Everything about how the session should run, set once at the start.
Messages two through twenty: normal work. You glance at the working state after each response. If a decision is wrong, you fix it right then.
Around message twenty: something feels soft. Run the drift check. Fix what drifted.
Around message twenty five: run the reset. Copy the memo into a fresh chat with the same three standing instructions. Keep going on a clean desk.
End of session: run the reset one more time, even if you're done. Save the memo and the attic. Tomorrow's session, or next month's, starts from that memo.
It sounds like overhead. In practice it's maybe two minutes of extra prompting per hour. And it fixes the main reason people give up on hard AI work: the session got dumb halfway through and they couldn't tell why.
When you're working with agents
Everything above applies double when you're running AI agents on longer jobs. Agents are exactly where the TraceML finding bites hardest, because nobody is sitting there noticing the loop.
Two adjustments for agent work.
Write the pivot rule into the agent's instructions. "If two attempts at the same approach fail, stop, write down why, and try an approach that starts from a different assumption. If three different approaches fail, stop and report what you tried." That last clause matters. An agent with no stop rule will happily burn tokens on attempt forty one.
Have the agent keep a running notes file. Same idea as the working state, but written to an actual file or document the agent can read back. Goal, constraints, what's been tried, what was learned. When the agent's context fills up and it starts over, it reads the notes first. That's the context management instruction at work, and it's pretty much what the Meta paper found.
If you're running agents through Make.com, the notes file can be a simple data store record that each run reads at the start and updates at the end. Ugly, effective, and it stops the same mistakes from repeating every run.
What this doesn't fix
Being honest about the limits.
These prompts make long sessions much more reliable. They don't make a model smarter than it is. If the problem is beyond the model, a clean desk just means it fails more clearly, which is still useful, because you'll know to bring in a person or a different tool.
They also depend on you reading the working state. The block is only as good as the attention you pay to it. If you scroll past it every time, you'll miss the moment a decision got recorded wrong.
And they're not worth the setup for quick tasks. Rewriting a paragraph, answering a question, drafting a short email. Just ask. Save this kit for the work that takes an hour or more, which, if you're doing it right, is where AI is actually earning its keep.
The habit underneath
The real lesson from both papers is the same one good managers learn the hard way.
Long, complex work fails when nobody's tracking the decisions, and it stalls when nobody's willing to say "this approach isn't working, let's try something different." That's true of teams of people and it turns out it's true of models too.
The upside with a model is you write the habit down once and it actually sticks. Keep the working state, clear the desk on a schedule, and make pivoting a rule so it doesn't depend on anybody's mood that day.
FROM THE AI NEWSROOM
The AI Business Accelerator
For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.
Join the Accelerator for $97
.....
This week
Pick one real piece of work you'd normally spend an hour or more on with AI. A proposal, a plan, a big email sequence, an analysis.
Start the session with the working state prompt, the pivot rule and the attic instruction together.
Work normally. When you hit twenty messages, run the reset and continue in a fresh chat from the memo.
At the end, save the final memo and the attic. Then compare the back half of that session to how your long ones usually go. I'm betting you won't want to go back.
Jordan
The AI Newsroom | Practical AI for people with a business to run.

