Harvey is a legal AI company valued at $15.6 billion. According to Bloomberg, its gross margin went from roughly positive 50 percent to negative 50 percent by June.
Negative fifty. For every dollar a customer paid Harvey, Harvey was paying its model providers about a dollar fifty.
It wasn't a pricing mistake in the usual sense. Customer token usage rose about twenty times under the usage based pricing Harvey was paying OpenAI and Anthropic. Customers got more value, used the product harder, and every bit of that extra usage came straight out of Harvey's margin because what they paid Harvey didn't move with what Harvey paid for tokens.
Harvey fixed it by launching an in house model in August, post trained on Moonshot's Kimi K3. Margins swung back positive.
You're not Harvey. You're probably not going to post train a model. But if you've built AI into anything you sell, whether that's a service, a done for you offer, a membership with an AI feature, or an agency retainer where you quietly use AI to deliver, you might have the same problem at a smaller scale and not know it yet.
Let's go find out.
The mismatch that creates the leak
Here's the shape of the problem in one sentence. You charge in one unit and you pay in another.
Your customer pays you monthly, or per project, or per seat. Flat. Predictable. That's good for them and good for your sales process.
You pay for AI per token, per call, per minute of video, per generation. Variable. It moves with how much work gets done.
As long as usage stays roughly where you guessed when you set your price, everything's fine. The trouble starts when usage grows, and here's the uncomfortable part: usage grows fastest when your product is working. The customer who gets the most value uses it the most, and the customer who uses it the most costs you the most. Your best customer can quietly become your least profitable one.
In services, this shows up differently but it's the same leak. You quoted a retainer based on how long the work took. Then you added an AI step that saves you hours, and then another, and now a chunk of your delivery runs on a paid API. Your hours went down and your costs went up in a new line item you've never actually matched against the client.
The margin audit
This takes about two hours and you'll know more about your business at the end of it than most of your competitors know about theirs.
Step one: list every place AI cost touches revenue. Go through your offers one at a time. For each one, write down every AI tool or API that gets used in delivering it. Not the tools you use for your own admin. The ones that scale with customer work. The transcription on client calls. The model that drafts client deliverables. The image generation for client creative. The chatbot on the client's site that you manage.
Step two: find the actual cost per customer. This is the step people skip because it's annoying. Pull last month's bill from each provider. Most of them let you break usage down by API key, project or workspace. If you haven't been separating usage by client, you'll have to estimate this month and fix the setup for next month.
The fix is simple: one API key per client, or one project per client, wherever your provider supports it. Tag it with the client name. From next month on, your bill tells you exactly who costs what.
Step three: rank your customers by AI cost as a share of what they pay you. Take each customer's monthly AI cost, divide by their monthly revenue, and sort the list. You're looking for the top of that list.
In most businesses I've seen do this, the distribution is lopsided. Most clients sit at a few percent. Two or three sit way higher, sometimes above twenty or thirty percent, and nobody noticed because those clients are happy, active and pay on time.
Step four: trend it. Do the same calculation for three months ago if you have the data. Is the share going up? For which clients? A client at 8 percent this month who was at 3 percent in June is telling you something about where they'll be in March.
Step five: ask what's driving it. For every client in the top of the list, find the specific workflow eating the budget. It's almost always one thing. One client asks for twelve revisions on every deliverable and each one's a full regeneration. One client's chatbot got popular. One client sends you ninety minute calls to transcribe and summarize every week instead of the thirty minutes you planned for.
FROM THE AI NEWSROOM
The AI Margin Kit
The client cost tracker, the five step margin audit with worked examples, the model re routing test sheet, the fair use tier template, and the pricing stress test I run before I put a price on anything with AI inside it. Everything in this issue, built out and ready to copy. Free for readers.
Reply MARGIN and I'll send it over
.....
The fixes, cheapest first
Once you know where the leak is, the fixes are mostly boring. That's good news.
Fix one: re route to a cheaper model. This is the one that got a lot easier last week. On September 22, Anthropic released Opus 5.5 at $4 per million input tokens and $20 output, which Anthropic says costs about 40 percent less to run than Opus 5 on typical work. OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50, both roughly half the price of the models they replace.
If your expensive client workflow is running on a flagship model because that's what you picked eighteen months ago, test it on a cheaper tier. For summarizing, tagging, extracting and first drafts, the cheaper tier is very often fine. Run five real client inputs through both and compare. If you can't tell the difference, neither can your client.
One caution: cheaper isn't always better on your specific task. OpenAI's own numbers showed Sol scoring lower than the model it replaces on some agentic coding benchmarks. Test on your real work before you switch.
Fix two: use caching. If a workflow sends the same long context over and over, like a client's brand guide, style rules or knowledge base, prompt caching can cut the cost of that repeated part dramatically. Opus 5.5 dropped its cache read price from fifty cents to twenty cents per million tokens. For agent loops and anything with a big reusable system prompt, the cache price matters more than the headline price, and most people have never turned it on.
Fix three: cap the revisions. If the leak is a client who regenerates everything repeatedly, the fix isn't technical. It's scope. "Two rounds of revisions included" is a sentence agencies have used for decades. It still works. AI made revisions feel free because they're fast. They aren't free, they just moved from your time to your API bill.
Fix four: add a usage tier. If some clients genuinely use far more than others, and they're getting more value, charge for it. A base retainer with a fair use level, plus a higher tier for heavy usage, is standard and nobody reasonable objects to it. The trick is setting the fair use level from the data you just pulled, not from a guess.
Fix five: pass it through. For some offers, the cleanest answer is to bill AI costs at cost, or at cost plus a small handling margin, as a separate line. Transparent, fair, and it removes the leak entirely. This works especially well for clients who are sophisticated about AI and would rather see the real number than pay a padded flat fee.
Pricing new offers so this never happens
If you're launching anything with AI inside it this quarter, build the math in from the start. Here's the checklist I use before I put a price on anything.
What does this cost me per customer at expected usage? What does it cost at three times expected usage? What does it cost at ten times? If the ten times number wipes out my margin, what's the mechanism that stops that from happening, whether that's a cap, a tier, a fair use policy or a pass through?
Most people only do the first calculation. Harvey's customers went to twenty times. Your heavy users won't get there, probably, but three to five times is completely normal for engaged customers once they figure out what the product does. Plan for it.
The other thing to plan for: prices change in both directions. Last week they went down, sharply. But introductory and promotional pricing has end dates. Some of the pricing that looks great this month has a printed expiration on it later this year. When you set a price to your customer, set it on the cost you'll be paying in six months, not the promotional number you're paying today.
What the big picture says about this
There's a reason both labs cut prices on the same day. The mid tier of the market filled up fast, with Grok 4.7, Xiaomi's MiMo V2.6 and StepFun's Step 5 all landing in the same week at a fraction of flagship prices. And Harvey showed everyone in public what happens to an application business when model costs scale faster than revenue.
That's good for you. The cost of AI capability is dropping faster than almost any input cost in any business you've ever run. But the savings don't show up in your margin on their own. You have to go get them, by re routing, by caching, by pricing correctly. The companies that do this quarterly will run at better margins than the ones that set their model once and forgot about it.
The broader lesson from Harvey isn't "build your own model." It's that the margin lives in the layer you control. For a big company that's the model weights. For you it's the routing, the prompts, the caching, the scoping and the price. All of those are things you can change this week.
If you don't sell anything with AI in it
Maybe you read this far and thought, I only use AI internally. It's not in my offers.
Run the audit anyway, pointed inward. List every AI subscription and API your business pays for. Add up the monthly total. Then, for each one, write down which revenue generating activity it actually supports.
You'll likely find two things. Some tools are carrying real weight and are obviously worth it. And some are subscriptions you started in a burst of enthusiasm, used for three weeks, and have been paying for ever since. If you'd like proof of which tools you actually use and for how long, Rize tracks where your hours actually go, which settles a lot of arguments with yourself.
Cancel the dead ones. Consolidate the overlapping ones. If you're paying for three separate chat subscriptions, a single aggregator login through Galaxy.ai might replace all three for less than you're paying for two of them.
FROM THE AI NEWSROOM
The AI Business Accelerator
For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.
Join the Accelerator for $97
.....
This week
Pull last month's invoice from every AI provider you pay for.
Pick your single biggest client or your single most popular offer. Estimate what it cost you in AI last month. Divide that by what it paid you.
If the number is under five percent, you're fine, go have a nice Tuesday. If it's over fifteen, you just found the most valuable hour of work you'll do this month.
Then create one API key per client going forward, so next month you don't have to estimate.
Jordan
The AI Newsroom | Practical AI for people with a business to run.

