Here's a sentence that would have sounded insane two years ago. In the last two weeks, at least five companies released or repriced a model that could plausibly be the right one for your business.
Anthropic shipped Claude Sonnet 5.5 last Monday at $2 per million input tokens and $10 output, with a million token context window and, by Anthropic's numbers, about 30 percent faster than the version it replaces. Google released Gemini 4 Argon, its first Gemini 4 frontier model, at an introductory $2 and $10 that rises to $4 and $20 later. Opus 5.5 topped the Epoch Capabilities Index this week. OpenAI's GPT-6 Sol and Luna arrived the week before. And the cheaper middle of the market kept filling up behind all of them.
So what's the right model for writing your sales emails now? For summarizing calls? For working through a pricing spreadsheet?
Honest answer: I don't know, and neither does anyone writing a benchmark chart. Benchmarks tell you how a model does on the benchmark. They don't tell you how it does on your clients, your voice, your messy inbox.
The only way to know is to test on your own work. And what kills most people's testing is plain old friction. Five accounts, five logins, five bills, five tabs, copying the same prompt five times.
This week's tool is Galaxy.ai, and I'm reviewing it now for one reason. This is the exact week it earns its money.
What it is
Galaxy.ai is an all in one AI subscription. Instead of paying separately for the major chat models, image generators, video tools and the rest, you pay one subscription and get access to a large catalog of them in one place.
That's the pitch. Whether it's worth it comes down to how many tools you'd otherwise pay for and how often you bounce between them.
I've reviewed it here before, back in the spring, so I'm not going to rehash the full feature tour. This one's narrower. I'll show you how I use it for my own model testing, then tell you where it comes up short.
The bake off routine
Here's the process I run whenever a big model release lands. It takes about an hour. More than once it's talked me out of chasing the shiny new thing. A couple of times it's pushed me to switch when habit would've kept me put.
Step one: pick three real tasks. Not test prompts you invent. Real work from the last two weeks that you'd actually do again. I use one writing task, one analysis task, and one task that requires following a long, detailed set of instructions. Those three cover most of what a business uses AI for.
Step two: freeze the inputs. Save the exact prompt and the exact source material for each task in a doc. Same words, same files, every model. If you tweak the prompt between models, you're testing your prompting, not the model.
Step three: run each task through four or five models. This is where having them in one place matters. Same prompt, paste, run, next model. Save every output into the doc under the model's name.
Step four: blind the outputs. Almost nobody bothers with this one, which is a shame, because it's the step that counts. Have someone else, or a simple script, strip the model names and shuffle the order. Label them A, B, C, D. You'd be amazed how much your opinion of a response changes when you don't know which famous lab wrote it.
Step five: score against a rubric you wrote before you looked. For the writing task: does it sound like me, is it the right length, did it include the specific details I asked for, would I send it with only light edits. For analysis: are the numbers right, did it catch the thing I know is in the data, did it make anything up. For instruction following: count the instructions it broke.
Step six: unblind and decide. Note the winner for each task type. Write it down somewhere your team can see it. "As of October, drafts go to model X, spreadsheet work goes to model Y."
That's a model routing policy for your business, built on your own work. You redo it whenever a major release lands, which lately means roughly monthly.
What I found this round
I ran the routine this week across the models available to me. I'm not going to crown a winner, because your three tasks aren't my three tasks and the whole point is that you run your own. But a few patterns are worth passing on.
The gaps are smaller than the marketing. On my writing task, three of the five outputs were good enough to send with light edits. The differences were in voice and length discipline, not in quality. If anybody tells you one model is dramatically better at everything, they're selling something.
Instruction following is where models actually separate. The task with fourteen specific requirements produced the biggest spread. One model nailed thirteen. Another hit nine and was confidently wrong about having done all fourteen. That's the test I'd weight most heavily for anything going into an automation.
Price per useful output beats price per token. A cheaper model that needs two rounds of revision costs you more than an expensive one that's right the first time, because your time is the expensive part. Count that when you pick.
The blind step changed my mind once. I would have bet money on which output I'd prefer for the analysis task. Blinded, I picked a different one. That alone justifies the routine.
FROM THE AI NEWSROOM
The AI Workflow Blueprint
The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.
Get the Blueprint for $47
.....
Where it genuinely saves money
Do the math on your own stack before you decide.
If you're paying for two or three separate AI chat subscriptions because different people on your team like different models, or because you use one for writing and another for research, a single bundle often comes in cheaper than the separate plans combined. List what you pay for now, add it up, and compare.
If you use image and video tools occasionally, maybe a few times a month for a social post or a thumbnail, a bundle beats paying monthly for tools you open twice. HeyGen is one I reviewed last week that sits in that category for a lot of people.
If you only use one model and you're happy with it, a bundle probably isn't for you. You'd be paying for a catalog you won't open.
Where it falls short
Now for the part the sales page skips.
You lose some native features. The individual labs keep building features into their own apps: projects with stored files, connectors to your email and drive, memory across conversations, agent modes, desktop tools. An aggregator gives you access to the model, but not always to everything the model's home app does. If a specific feature is central to how you work, check whether the bundle supports it before you cancel the direct subscription.
New models can take time to show up. When a lab releases something, it's available in that lab's own app on day one. How quickly it reaches any aggregator varies. For a release week like this one, check the model list before you assume the newest version is there.
Usage limits work differently. Bundles have their own limits and they don't always map one to one with what you'd get on a direct plan. If you're a very heavy user of one specific model, compare the limits carefully.
API access is separate. If your automations call models through APIs, the bundle doesn't replace those. Your Make.com scenarios still need their own API connections. The bundle is for the human side of the work: testing, drafting, research. Which, to be fair, is where the bake off happens.
How I actually use it
My setup, for what it's worth.
I keep direct subscriptions to the one or two apps whose native features I live in every day. Projects, connectors, the desktop tools. That's my main workbench.
I use the bundle for everything around the edges. Model testing whenever a release lands. Second opinions, where I want a different lab's model to check a first draft for gaps. Occasional image and video work. Trying a new tool before deciding whether it deserves its own subscription.
That split gets me the deep features where I need them and breadth everywhere else, for less than I was paying before I consolidated.
Who should buy it
Buy it if you pay for two or more AI subscriptions today and use them for overlapping work. The consolidation math probably works.
Buy it if you want to run your own model testing and the setup friction has stopped you. That's the best reason to grab it this month.
Buy it if you're a small team where different people prefer different models and you'd rather manage one bill.
Skip it if you live inside one lab's app, use its native features heavily, and never feel the need to compare. Your money's better spent on the higher tier of that one app.
Skip it if your AI use is entirely through APIs and automations. You need API accounts, not a chat bundle.
The bigger point
Model fatigue is real. CNBC ran a whole piece on it last month. Releases are coming so fast that a lot of business owners have just stopped paying attention.
I understand the impulse, but tuning out completely is a mistake. The price and quality differences between models are big enough now to matter to your margins and your output. What doesn't work is chasing every release on vibes.
The middle path is a routine. Same three tasks, same rubric, blind scoring, an hour each time something big lands. You stay current without getting dragged around by every launch announcement, and you end up with a written routing policy that's built on your work instead of somebody's leaderboard.
Use whatever makes the routine easy enough that you'll actually run it. Mine's a single tab with all the models in it.
FROM THE AI NEWSROOM
The AI Business Accelerator
For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.
Join the Accelerator for $97
.....
This week
Pick your three tasks. One writing, one analysis, one long instruction list. Real work from the last two weeks.
Freeze the prompts and inputs in a doc. Run each one through at least three different models, ideally including one of this week's new releases.
Blind the outputs, score them against a rubric you wrote first, then unblind.
Write the winners down as your routing policy for October. Put a reminder on the calendar to run it again the next time a major model ships, which, at this pace, will be about a week from Tuesday.
Jordan
The AI Newsroom | Practical AI for people with a business to run.

