THE PROMPT VAULT

The failure mode nobody warns you about isn't the model being wrong. It's the model being wrong in a tone that sounds exactly like being right.

A wrong answer that sounds uncertain is nearly harmless. You go check it. A wrong answer delivered with the same steady confidence as a correct one is how bad numbers end up in a board deck and how a pricing decision gets made on a competitor analysis that was partly invented.

This isn't a reason to stop using the things. It's a reason to change how you ask.

Every prompt below does the same job in a different context: it makes the model expose the shape of its own reasoning so you can see where to push. None of them are clever. They're the seven I actually run, and they've caught enough real errors that I no longer work without them.

Copy them, adjust the bracketed parts, keep them somewhere you can reach fast.

One. The confidence split

The problem with a normal answer is that everything arrives at the same volume. The parts the model is sure about and the parts it's reconstructing from vibes look identical.

Answer the question below. Then split your answer into three lists. First, the claims you are confident about and why. Second, the claims that are probably right but that you would want verified before anyone acted on them. Third, anything you filled in because it seemed reasonable rather than because you actually knew it. Be honest about the third list. If it's empty, say so and explain why you're sure it's empty.

The question: [your question]

The third list is the whole point. Models will populate it if you make it a normal, expected part of the output rather than an admission of failure. What you get back is an answer with the risky parts already flagged, which turns a review from reading everything sceptically into checking three specific things.

Two. The source of the number

Numbers are where this matters most, because a number carries an authority that a sentence doesn't. People argue with prose and accept figures.

For every number in your answer, tell me where it came from. Choose one label for each: a figure I gave you in this conversation, a figure you calculated from what I gave you, a widely published figure you're recalling, or an estimate you produced. For calculated numbers, show the arithmetic. For estimates, give me the range you'd actually defend rather than a single number, and say what would move it.

The label "estimate you produced" is the one that saves you. There is a real difference between a market size you pulled from a report and a market size assembled from plausible reasoning, and by default both arrive looking like facts.

The range request does something useful too. A model asked for one number gives you a confident one. Asked for a range it will defend, it usually gives you something honest and wide, which tells you immediately how much weight the figure can carry.

Three. The steelman before the answer

For any decision, the danger is that you've already half decided and you're using the model to feel better about it. Models are agreeable. If you lean, they lean with you.

Before you tell me what you think, argue the strongest possible case for each option. Not a balanced summary, the actual best case, the one a smart person who genuinely believed it would make. Then tell me which case is stronger and, specifically, what fact would have to be true for the other one to win instead. End with what you would need to know that you currently don't.

That last clause is the useful one. "What you would need to know that you currently don't" converts an opinion into a research list, and the research list is what actually improves the decision.

Four. The assumption audit

This one is for anything where the model is working from a description you wrote, which is most business use.

List every assumption you made to produce that answer, including ones you think are obvious. For each, mark whether it came from something I actually told you or from something you inferred. Then tell me which two assumptions, if wrong, would most change your answer.

The gap between what you told it and what it inferred is where most bad AI output lives. You describe your business in four sentences, the model fills in forty more from the average business it has seen, and then it gives you advice for that average business.

Seeing the inferred list is frequently uncomfortable. It's also the fastest way to discover that you under briefed it.

THE AI WORKFLOW BLUEPRINT • $47

The Blueprint includes the full prompt library with the verification patterns built in, plus the context templates that stop the model inferring your business wrong in the first place. Better inputs beat better prompts, and the templates are how you get better inputs consistently.

Five. The second pass

Never trust a first draft, including one you asked for well. This prompt goes after any output that matters.

Review the answer you just gave me as if a competitor wrote it and you were being paid to find its weakest point. What is the single most likely thing to be wrong? What did it skip that a careful person would have included? Where is it confident in a way the evidence doesn't support? Do not be polite about it.

Running a model against its own output works better than it has any right to. It's a different task, critiquing rather than generating, and it produces genuinely different results. The instruction to skip the politeness matters, because the default behaviour is to find three minor things and call it a review.

Six. The specifics extractor

This is the one I use most, and it's less about verification than about forcing usefulness.

That answer is too general to act on. Rewrite it so that every recommendation includes who does it, what the first concrete step is, roughly how long that step takes, and how I would know within two weeks whether it's working. If any recommendation can't survive that treatment, delete it rather than padding it out.

The last sentence is doing the work. Most generic advice cannot survive being made specific, and telling the model to delete rather than pad prevents it from dressing up a vague suggestion in fake precision. What comes back is shorter and usable.

Seven. The refusal check

For anything where you're asking a model to work with your own data, particularly numbers and reports.

If you don't have enough information in what I've given you to answer this properly, say so and tell me exactly what's missing. Do not produce a best effort answer with the gaps filled in. I would rather have a short list of what you need than a complete looking answer built on holes.

Models default to being helpful, and being helpful reads as producing something. Explicitly giving permission to come back empty handed changes the behaviour, and it's the difference between finding out you're missing a data source now versus after a decision.

The prompt that fixes the other seven

There's a failure the checks above will surface but not solve, and it's worth addressing directly because it's the most common one.

Most bad output is not a prompting problem. It's a briefing problem. You asked a good question with four sentences of context about a business the model has never seen, and it did the only thing it could, which was assume you resemble the average case.

The assumption audit catches this after the fact. This catches it before.

Before you answer, ask me up to six questions about my situation that would most change your answer. Ask only about things you genuinely don't know and that actually matter to the recommendation. Rank them so the most important is first. Do not answer the question yet.

Then answer them, then ask the real question.

The reason this works is that it inverts who's responsible for identifying the gaps. You don't know what the model needs, because you don't know what it's assuming. It does know, or at least it knows what varies between cases, and asked directly it will tell you.

The instruction to only ask about things that matter is load bearing. Without it you get six generic intake questions of the kind a bad consultant asks to look thorough. With it you usually get two or three genuinely sharp ones, and answering them takes two minutes and improves the output more than any amount of prompt engineering on the question itself.

Run this any time the stakes are above trivial. The two minutes it costs is the cheapest quality improvement available.

How to actually use these

A few things I've learned running these in real work.

Don't stack them into one giant prompt. It's tempting and it produces worse results than running them in sequence. Each one is a distinct task, and asking for six distinct tasks at once gets you six shallow attempts.

Run one, two and six on almost everything. Those three catch the majority of problems for very little effort. Bring in three and four for decisions, five for anything going in front of another human, seven for anything involving your own data.

Save them where you'll use them. A prompt library you have to go find is a prompt library you don't use. Wherever your snippets live, put these there.

And watch what the checks turn up over the first fortnight. If the confidence split keeps flagging the same category, that's not a prompting problem, that's a signal you're asking the model to do something it isn't good at. Which is worth knowing early, and is genuinely the most valuable output of the whole exercise.

One more thing about how to read the results, because people get this backwards.

When a check flags something, the useful response is not to ask the model to fix it. It's to go and find out. If the source of the number prompt tells you a figure was an estimate, asking the model for a better estimate gets you a more confident sounding estimate, not a better one. Go find the actual number, or accept that you're working with a range and decide accordingly.

The prompts are instruments, not repairs. They tell you where the soft ground is. Walking across it anyway because the model reassured you on the second attempt is how you end up exactly where you'd have been without running the check, except now you feel diligent about it.

The habit worth building is small: run one, read the flags, and for each flag make an explicit decision to either verify it or proceed knowing it's soft. Written down, ideally, next to the output. Three months later when somebody asks where a figure came from, that note is worth more than the answer was.

The point of all this isn't scepticism for its own sake. It's that a model with its uncertainty visible is a much better collaborator than one performing certainty, and you get the visible version by asking for it.

THE AI BUSINESS ACCELERATOR • $97

Eight weeks of building AI into your business in a way that holds up. We set up the prompt library, work out which tasks in your business the models are actually reliable at, and put the verification habits in place so nobody has to remember to be careful.

Jordan

The AI Newsroom | Jordan Hale | ainewsroomdaily.com

Keep reading