Consulting Work Blog Contact
← Back to blog

I Cut My AI Bill by ~75% Without Losing Quality. Here's the Decision That Did It.

I Cut My AI Bill by ~75% Without Losing Quality. Here's the Decision That Did It.

The fastest way to overspend on AI is to use the most powerful, most expensive option for everything, including the boring 80% of work a far cheaper option handles just as well. That’s exactly what I was doing. Then I cut my monthly AI bill by roughly three quarters and, crucially, checked that quality held. No magic, no clever hack, just a handful of unglamorous decisions and one test that kept me honest. If you’re running AI across a team and the invoice keeps creeping up, this is the part worth copying.

TL;DR

  • Most of the cost was premium tooling doing work that didn’t need premium. I split the work into “needs the expensive option” and “doesn’t,” then measured the line.
  • Cheaper option for the routine bulk, premium reserved for the hard cases. Bill dropped ~75%; quality, checked on real data, stayed put.
  • The risk isn’t cutting cost. It’s cutting cost without measuring. That’s how quality quietly slips.

Monthly AI bill before and after, quality held

Where the money was actually going

When I looked, the picture was almost embarrassing. The overwhelming majority of my AI usage was simple, repetitive, low-stakes work, the kind where the answer is rarely subtle. All of it was running through the most capable, most expensive option, a frontier model like Claude Opus or GPT-4o, because that’s what I’d set up on day one and never revisited. That’s the most common AI overspend there is, and it hides in plain sight because the tool works. It’s just overkill.

The decision that did the work

I sorted the work into two buckets:

  • The routine bulk. High volume, low stakes, rarely ambiguous. The cheap option is genuinely good enough here.
  • The hard cases. Lower volume, higher stakes, where being wrong is expensive and the answer is often subtle. Worth the premium.

Then I moved the routine bulk (the large majority) onto a much cheaper option, a smaller, lighter model like Claude Haiku or GPT-4o mini, and kept the premium one for the hard cases only. That’s the whole move. The bill fell by about three quarters because the expensive tool went from handling everything to handling the small slice that actually justified it. The drop is so steep because of the price gap between tiers: a frontier model can cost more than ten times as much per unit of work as a lighter one.

The step that kept it honest

Here’s the part most cost-cutting skips, and it’s the part that matters. Cutting cost is easy. Cutting cost without quietly lowering quality is the actual skill, and you can’t eyeball it. A cheaper option will often look fine on the obvious cases and fall apart on the hard ones, which is exactly the failure you won’t notice until a customer or a colleague does.

So before committing, I tested the cheaper option on a slice of real, messy work and compared it against what I had. On the routine bulk it held up. On the hard cases it didn’t, which is precisely why those stayed on the premium tool. The split wasn’t a guess; it was drawn where the measurement said to draw it. (If you want the long version of why your own test set lies to you about this, that’s a whole separate story.)

A simple way to find your own savings

You don’t need to be technical to do any of this. Three questions cover it:

  1. What am I using the expensive option for? List it. Most people have never actually looked.
  2. How much of that is routine and low-stakes? That’s your savings pool, usually most of it.
  3. Does a cheaper option hold up on that routine work, tested on real data? If yes, move it. If no, you’ve found a case that genuinely needs the premium.

The savings live in the gap between “what we pay for” and “what the work actually requires.” For most teams that gap is large, and nobody’s looked at it because the expensive thing was working. It keeps working right up until you check the bill.

The savings aren’t in using less AI. They’re in stopping the expensive AI from doing cheap work.

What I’d do differently

I should have revisited this far sooner. I set up the expensive default once and left it running for months, treating the bill as fixed. It wasn’t fixed: it was a decision I’d made once and never re-examined, which is the most expensive kind of decision there is. Now I treat “which tool for which work” as something to review on a schedule. The options keep getting cheaper and better, and a default you chose a year ago is almost never the right one today.

The headline is the 75%, but the lesson underneath is duller and more valuable: most AI overspend isn’t from using AI too much, it’s from using the expensive AI for work that never needed it. The exact number depends on your mix, but the shape is common, and most teams are more lopsided than they think.

So here’s the question for your own invoice: how much of your premium AI spend is doing work a cheaper option could handle just as well, and when did you last actually test that?

Stack: Claude Opus · Claude Haiku · GPT-4o mini · Python

Need something like this for your own business? See how I can help →