Count the questions your automations ask a language model every day.

Not the big ones. Not "write a proposal" or "summarize this call." The small ones. Is this email a sales lead or a vendor pitch? Is this support ticket urgent? Does this invoice match the purchase order? Should this lead go to the calendar link or the nurture sequence? Is this comment spam?

In most businesses I look at, those little yes or no questions make up the majority of model calls. And almost all of them are being answered by the same big, expensive, slow model that writes the proposals.

That's like hiring a senior attorney to sort your mail.

Last week, quietly, the market started selling you a mail sorter. In the span of a few days, at least half a dozen companies shipped what they're calling decision models. Small, fast models built to do one thing: take an input and return a typed answer. Yes or no. One option from a list. A score. With a confidence number attached.

Perplexity launched a Decisions API priced at four cents per million input tokens. Cloudflare released open decision models called Clef and Clef flash under an Apache licence, with median latency around 39 milliseconds. Liquid put a model called D1 on OpenRouter at four cents in and nothing for output. Inception Labs shipped Mercury Decide. Strands released a two billion parameter model built for routing and tool selection. Ollama added support for typed yes, no, choice and score outputs on local models.

Six companies don't ship the same thing in the same week by accident. Somebody noticed how much money's getting burned on easy questions.

This one's about getting your small questions off your big model.

Why this matters more than it sounds

Three reasons, and the first one isn't money.

Speed. A big model answering a yes or no question still has to warm up, read your whole prompt, and generate a response. That can take a few seconds. A decision model built for the job comes back in tens of milliseconds. If you've got a workflow with five classification steps chained together, that's the difference between a process that finishes before the customer refreshes the page and one that doesn't.

Consistency. Ask a big model "is this urgent?" and you'll get "Yes, this appears urgent because..." one time and "URGENT" the next and a thoughtful paragraph the time after that. Then your automation has to parse all three. Decision models return a typed answer by design. Yes means yes. The field always has one of the values you allowed. Nothing to parse, nothing to break.

Cost. Yes, the money matters too. At four cents per million input tokens, you could classify every email your business will receive this decade for less than a sandwich. Even if you're only saving a few dollars a month today, the bigger point is that cheap decisions change what's worth automating at all. Things you skipped because a flagship call per item didn't pencil out suddenly do.

Step one: find your decisions

Open your automation platform and go scenario by scenario. You're hunting for any model call that spits out a label, a route or a flag the next step uses to decide what happens. No human ever reads the output. A machine acts on it.

Here's what usually turns up:

Routing. Which inbox, which team member, which pipeline stage, which follow up sequence.

Triage. Urgent or not. Needs a human or doesn't. Refund request or general question.

Qualification. Does this lead fit the ideal customer profile. Is this company over a certain size. Is this inquiry about a service you actually sell.

Validation. Does this document contain a signature. Does this invoice total match the order. Did the form submitter give a real business name.

Filtering. Spam or not. On topic or not. Duplicate or new.

Put every one of these in a list. One row per decision, with three columns: what the question is, what the allowed answers are, and what happens downstream for each answer.

That third column is the important one. If you can't say what happens next for each possible answer, you don't have a decision yet. You have a vibe you were hoping the model would turn into a process.

Step two: rewrite each one as a typed question

This is where most of the value actually comes from, and you can do it before you switch a single model.

Big models let you get away with sloppy questions. "Look at this email and tell me what to do with it" works, sort of, because the model is smart enough to guess. Decision models won't guess for you. They need the question shaped properly. And it turns out that shaping it properly makes your big model better at the job too.

Every decision gets rewritten into one of three shapes.

Yes or no. "Is this message a request for a refund?" Not "what kind of message is this." One question, one boolean.

Pick one. "Which of these best describes the message: sales inquiry, support request, billing question, vendor pitch, other." Always include "other." Always. Without it, the model has to jam a weird message into a category that doesn't fit, and you'll never find out.

Score. "On a scale of 1 to 5, how closely does this company match the following profile?" Then spell out what a 1 looks like and what a 5 looks like. A score with no anchors is a random number generator with extra steps.

Here's the template I use to rewrite them:

❝

You are making one decision. Read the input and answer only the question below.

QUESTION: [one sentence]

ALLOWED ANSWERS: [the exact list, including "other" or "unsure" where it makes sense]

DEFINITIONS: [one line per answer saying exactly what qualifies]

Return only the answer and a confidence between 0 and 1.

Those definitions are where you'll find out how fuzzy your own process is. Writing "urgent means the customer can't use the product right now, or a payment failed, or they've mentioned cancelling" forces you to decide what urgent means in your business. Your team has been guessing about that for years. Now it's written down.

Step three: use the confidence number

Cheap is nice. This is the part that makes them actually useful.

Most of these models return a probability or confidence alongside the answer. That number is your routing key. You're going to build three lanes.

High confidence lane. Above a threshold you set, say 0.9, the answer goes straight through. The automation acts on it. Nobody looks.

Middle lane. Between, say, 0.6 and 0.9, the decision gets escalated to your bigger model with the full context and a prompt that asks it to think it through. You're spending the expensive call only on the cases that need it.

Low confidence lane. Below 0.6, a human sees it. It lands in a review queue, a Slack channel, a sheet row, wherever you'll actually look at it.

In Make.com this is a router with three routes and a filter on each one. Five minutes to build. If you'd like the AI step to live right inside your CRM instead, the same three lane logic works in a Go High Level workflow with if and else branches on the confidence field.

Here's what usually happens when people turn this on. Somewhere between 70 and 90 percent of decisions land in the fast lane. A small slice goes to the big model. A trickle goes to a human. And the trickle is the most valuable part, because those are the weird cases your old setup was quietly getting wrong.

Step four: test it against what actually happened

Don't switch anything on trust. You've got history. Use it.

Pull fifty real examples of each decision from the last month, along with what the right answer turned out to be. If your inbox router sent a vendor pitch to the sales team, you know the right answer was "vendor pitch." Build a little test sheet: input, correct answer.

Run all fifty through the decision model. Run them through whatever you're using today. Compare.

You're looking at three numbers. How often each one got it right. How often the decision model was confident and wrong, which is the dangerous one. And how many it sent to the middle and low lanes.

If the decision model matches or beats your current setup on accuracy, and the confident and wrong count is near zero, switch. If it's close but the confident errors bother you, raise the threshold for the fast lane and let more go to the middle. If it's clearly worse, keep the big model for that decision. Not every question belongs on a small model, and you'll know which ones from the test, not from the marketing page.

Comparing a few models side by side is a lot less painful when you can try them without setting up separate billing for each one. I do my side by side testing through Galaxy.ai, then wire the winner into production.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

Where decision models fall down

They're not magic, and the week they launched is exactly the week to be clear eyed about it.

They don't explain themselves. You get an answer and a number. If you need a reason written into the record for compliance or for the customer, you need a bigger model or a human for that step.

They don't handle long context well. A decision model is great at reading an email. It's not the tool for reading a forty page contract and deciding whether clause nine creates a liability problem. Keep the long documents on the big models.

Confidence can lie. A model can be very confident and very wrong, especially on inputs that look nothing like what it saw in training. That's why the test against history matters, and why you should keep sampling the fast lane even after you trust it. I pull ten random fast lane decisions every Friday and check them by hand. It takes four minutes and it's caught drift twice.

Your categories might be the problem. If accuracy is bad across every model you try, the issue is usually the question, not the model. Overlapping categories, missing "other," definitions that leave room for argument. Go back to step two.

A real build, start to finish

Let me make this concrete with the one I'd build first in almost any service business: the inbound triage router.

Every message that hits your general inbox or contact form runs through it.

Decision one, pick one: sales inquiry, existing client request, billing, vendor or partnership pitch, job application, spam, other. Confidence threshold for the fast lane: 0.9.

Decision two, only for sales inquiries, score 1 to 5: how well does this match your ideal customer, with anchors written out for 1, 3 and 5. Four and above goes straight to a booking link reply drafted for your review. Two and three go to a nurture sequence. One gets a polite "not the right fit" draft.

Decision three, only for client requests, yes or no: does this need a response within four business hours? Yes routes to whoever's on point today with a notification. No goes into the normal queue.

Every decision writes three fields back to the record: the answer, the confidence, and the model that made the call. That last one matters because you'll switch models again in three months, and you'll want to know which era a decision came from.

Total build time, maybe ninety minutes including testing. Running cost at current decision model pricing, close enough to zero that you won't find it on the invoice. Here's what you'll actually notice. Nothing important sits in the general inbox for six hours just because the person who checks it was stuck in meetings.

If your messages come in through calls rather than email, the same logic works on transcripts. Fathom gives you the transcript, a decision model tags the call type and whether a follow up was promised, and the router does the rest.

The bigger shift

Forget the savings for a second. Here's what I think actually changed.

For two years, the default way to add AI to a process was to point the smartest available model at the problem and write a long prompt. That worked, but it trained a lot of people to think of AI as one big brain you consult about everything.

Decision models nudge you into a better habit. Chop the work into small questions and hand each one to the cheapest tool that gets it right. Save the pricey thinking for the parts that need it, and let a human take whatever the machine isn't sure about.

That's just good process design. It's how you'd staff a team if you were doing it well. You don't have the founder answer the phones. You don't have the intern sign contracts.

Rebuild your automations this way over the next quarter and you'll be quicker and cheaper than the folks still feeding every question to the flagship. You'll also catch the mistakes that used to slip by without a sound. And you'll get a bonus nobody talks about: a written list of every decision the business makes and exactly what each answer means. Most companies have never had that. It's worth having even if you never touch a model.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

This week

Pick one scenario. Just one, the one with the most volume.

Find every model call in it that returns a label instead of text. Rewrite each one using the typed question template, with real definitions for every allowed answer.

Pull fifty real examples from last month. Run them through a decision model and through your current setup. Count the right answers and the confident wrong answers.

If the small model holds up, add the three lane router and switch it on. Then set a reminder to check ten random fast lane decisions next Friday.

That's an afternoon. After that, your expensive model can stop sorting mail and get back to the hard stuff.

Jordan

The AI Newsroom | Practical AI for people with a business to run.