At 00:57 UTC on September 22, Anthropic opened an incident covering most of its top models. Fable and Mythos came back that morning. Opus 5 kept throwing elevated errors into a second day.

Opus 5 is the default model in a huge share of production pipelines right now. Coding agents, support triage, document processing, lead scoring. For a lot of businesses, an Opus outage wasn't an inconvenience. It was the whole machine sitting still while the invoices kept arriving.

And it wasn't a one off. It was the third capacity signal from Anthropic in fourteen days, after a cut to Claude Code weekly limits. OpenAI closed ChatGPT Pro to new signups on September 11 for the same reason. Demand is outrunning compute at every lab, and when that happens, somebody's requests get dropped.

Here's the question I want you to answer honestly before you read any further. If your main model went dark for thirty six hours starting tonight, what happens to your business?

If the answer is "I'd find out when a customer emailed me," this one's for you.

Single points of failure you didn't choose

When you built your first automations, you picked a model. Probably the best one available at the time, because why wouldn't you. You wired it into a scenario, tested it, and moved on.

Then you built the next one. Same model, because it worked. And the next. Eighteen months later you've got twenty workflows, and every one of them depends on a single vendor's single model being up.

Nobody designed that. It accreted. And it's exactly the architecture that a two day outage exposes.

Here's the thing people miss. Model outages don't look like outages from inside your automation platform. The scenario runs. The HTTP module fires. It gets back a 529, or a 503, or a timeout after ninety seconds. Depending on how you set it up, that either stops the scenario with an error nobody reads, or it quietly passes an empty value downstream and your system sends a blank email to a customer.

The second one is worse, and it's more common than you'd think.

The failover build, step by step

I'm going to walk through this in Make.com because that's where I build, but the logic is identical in any platform with error handling. It takes about an hour for your first workflow and fifteen minutes for each one after that.

Step one: wrap every model call in an error handler. In Make, right click the module that calls the model and add an error handler route. If you don't have one, a failed call either stops the whole scenario or, with some settings, gets ignored. You want neither. You want a decision.

Step two: retry once, briefly. A lot of failures during capacity crunches are transient. The first request gets a 529 overloaded response, and the same request thirty seconds later goes through fine. Add a short sleep, then retry the identical call. One retry. Not five. Hammering an overloaded endpoint five times in a row doesn't help you and it doesn't help anybody else either.

Step three: fail over to a different vendor. If the retry fails, route to a second model from a different company. Not a different model from the same company, because when a lab has a capacity problem it usually affects more than one model at a time. That's exactly what happened on the 22nd, when the incident covered Fable, Mythos and Opus together.

Step four: mark what happened. Whatever model produced the output, write its name into a field that travels with the record. If a customer reply went out from your fallback, you want to know that later. This is one extra variable and it saves you hours of detective work.

Step five: queue what can't fail over. Some tasks shouldn't run on a fallback. Maybe the output quality matters too much, or the prompt is tuned so tightly to one model that another one produces junk. For those, don't fail over. Park the input in a queue, a data store or a sheet row with a status of "waiting," and have a separate scenario reprocess the queue every hour once the primary model responds again.

That's it. Five steps. Retry, fail over, label, and queue whatever shouldn't fail over. The whole thing is maybe six modules.

Picking the fallback model

This used to be the hard part. It isn't anymore, because the market just handed you a dozen good options at a fraction of flagship prices.

In the forty eight hours around September 21, three models landed in the same capability band. xAI shipped Grok 4.7 at $2 per million input tokens and $6 output. Xiaomi released MiMo V2.6 under an MIT licence, with its Pro model scoring 46 on the Artificial Analysis Intelligence Index. StepFun put out Step 5 Preview. Then on the 22nd, OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at ten cents in and fifty cents out, and Anthropic answered with Opus 5.5 at $4 and $20.

For a fallback, you don't need the best model on the market. You need a model that produces acceptable output for that specific task, from a different vendor than your primary, at a price you won't wince at when it carries your whole load for a day.

Here's how I'd pick, by task type:

Classification and extraction. Sorting inbound email, pulling fields from invoices, tagging leads. Almost any current model handles this. Pick the cheapest one from a different vendor. For a lot of people that's going to be Luna now, which at ten cents per million input tokens is cheap enough that the fallback bill during an outage is a rounding error.

Drafting customer facing text. Replies, follow ups, summaries a human will read. Pick a mid tier model from a different vendor and test it against five real inputs from last week. If three of the five would have been fine to send, it's a good enough fallback. If not, this task goes in the queue instead.

Anything agentic. Multi step work where the model calls tools and makes decisions. Be careful here. Agent behavior varies a lot between models even when the benchmark scores look similar. My default is to queue agentic work rather than fail it over, unless I've specifically tested the fallback on that exact workflow.

If juggling four API accounts sounds like a headache, that's a fair objection. An aggregator like Galaxy.ai gives you a pile of models under one login, which makes the testing phase a lot faster. You can throw the same five inputs at four models in an afternoon without setting up billing four times.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

The part that breaks when you switch models

Here's where failover goes wrong for most people. They build the routing, they test it by forcing an error, the fallback fires, and they call it done.

Then the real outage comes and the fallback produces output in a slightly different shape. The primary model returned clean JSON. The fallback returns clean JSON wrapped in a friendly sentence. Your parser chokes, every record after that point fails, and now you've got an outage from a model that's actually working fine.

Three fixes for this, in order of importance.

Use structured output everywhere you can. Most current models support a mode where you hand over a schema and the model is forced to match it. If both your primary and your fallback support it, use it on both, and shape differences mostly go away.

Put a validation step after every model call. Before the output goes anywhere, check it. Does the category field contain one of your allowed values? Is the summary under a length limit? Is the email field an email? If validation fails, treat it exactly like a failed call and route to the next option. This single module catches more problems than anything else in the build, including problems that have nothing to do with outages.

Keep a separate prompt for the fallback. Prompts get tuned to a model over time, whether you meant to or not. The phrasing that works beautifully on one model can confuse another. Keep a fallback version of each prompt, test it on the fallback model, and store both. It's a little more maintenance and it's the difference between a fallback that works and one that just technically runs.

Test it like you mean it

You can't know if your failover works until you watch it work. And you don't want the first time you watch it to be during an actual outage at two in the afternoon on a Tuesday.

Once a month, deliberately break your primary. The easy way: swap the API key in your primary connection for an invalid one, or point the model name at something that doesn't exist. Run five real inputs through. Watch what happens.

You're checking four things. Did the retry fire? Did the fallback fire? Did the output pass validation and land where it's supposed to land? Did the "which model produced this" field get written correctly?

Then put the real key back. Total time, about fifteen minutes. I do it on the first Monday of the month and I've found a broken failover path three times this year. Every single one would have failed silently during a real incident.

What to do during an actual outage

Having the routing built is most of the battle. But there are a few things worth doing by hand when you know a provider is having a bad day.

Watch the status page, not social media. Every lab has one. Subscribe to email or text alerts on the status page for every model provider you depend on. It takes two minutes, and you'll know about problems before your customers do. That's the whole point.

Pause anything that's queue only. If a workflow is set to queue rather than fail over, and the outage is looking like a long one, check how big the queue is getting. A queue of forty items reprocessing when the model comes back is fine. A queue of four thousand might hit rate limits on recovery and cause a second, smaller mess. For big queues, reprocess in batches.

Tell the humans. If your team relies on AI assisted tools, a short message saying "the main model is having issues today, expect slower responses, here's the workaround" prevents a lot of confused tickets and at least one person trying to fix something that isn't broken on your end.

Write down what happened. After it's over, spend ten minutes noting what failed over, what queued, what broke anyway, and what you'd change. That note is worth more than any vendor's post mortem because it's about your system, not theirs.

The cost math that makes this obvious

Some people resist this because it feels like paying for insurance they won't use. So let's do the math.

Setup time: an hour for the first workflow, fifteen minutes per workflow after that. For a business with twelve important automations, that's about four hours of build time total.

Ongoing cost: basically zero when nothing's broken, because the fallback model only gets called when the primary fails. The monthly test costs a few cents in tokens.

Now the other side. Last week's incident lasted into a second day for Opus 5. If you process inbound leads with AI, what's the value of a day and a half of leads that got no response, or a delayed one? If you run support triage through a model, what does thirty six hours of backed up tickets cost you in customer patience?

For most businesses with any real volume, one outage costs more than the four hours of build time. And there will be more outages. Anthropic passed a $100 billion run rate this month on infrastructure that was already tight enough to trim limits. Demand isn't slowing down, and new capacity takes years to build. Plan accordingly.

One more reason to build this now

There's a second payoff here that has nothing to do with outages.

Once you've built failover, you've built the ability to swap models with a configuration change. That matters a lot in a market where the price of comparable capability just dropped by half in a single day. Opus 5.5 came in about 40 percent cheaper to run than Opus 5. Sol and Luna came in at roughly half the price of the models they replace.

If your workflows are hard wired to one model, taking advantage of a price cut means editing every scenario by hand. If they route through a primary and a fallback that you control from one place, you change one setting and every workflow picks it up.

So the failover work isn't just insurance. It's the plumbing that lets you move whenever the pricing moves, which lately is every couple of weeks.

I keep the whole setup documented in a Notion page, one row per workflow, with the primary model, the fallback model, whether the task fails over or queues, and the date of the last test. It's the least exciting document in my business and I'd rescue it from a fire before almost anything else.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

This week

Open your automation platform. Find the one workflow that would hurt most if it stopped for two days.

Add an error handler to its model call. Add one retry. Add a fallback to a model from a different company. Add a validation check.

Then break your primary on purpose and watch the fallback fire.

That's one workflow, about an hour, and the next time a lab has a bad Tuesday, you'll be the person in your group chat who didn't notice.

Jordan

The AI Newsroom | Practical AI for people with a business to run.