C:\CHANGELOG> type v1-33-0-the-bill-went-down-the-fleet-did-not.md
v1.33.0 · released · 5 min read · by

The bill went down. The fleet did not.

Ramp dropped its September AI Index three days ago. Anthropic hit 43.8% of U.S. businesses, up 0.34 points month over month. OpenAI reached 39.8%, up 0.09 points. Adoption is still climbing. The interesting number is the one that went the other direction: AI spending per employee among leading spenders declined. Not stalled — declined. Companies that were burning the most tokens are now burning fewer of them, on purpose, while keeping the same systems running.

I run ten production platforms. The ERP at coenconstruction.com has been calling models since GPT-3.5-turbo cost five times what it does now. The estimating SaaS talks to Anthropic. The review-request dialer on valhalla-k9.com and hiremariaelena.com runs on Twilio and hits Claude for message rewriting. Every one of those integrations used to cost more than it does today — not because I rewrote them, but because the same API call got cheaper three times in eighteen months. The infrastructure I built didn't change. The bill did.

Ramp's data shows something different. Customers are choosing cheaper models and eliminating unnecessary calls, which means they're not riding the price cuts passively — they're actively re-architecting to spend less. That's not a sign the models got worse. It's a sign businesses learned what actually needed the frontier model and what could run on the smaller one at a tenth of the cost. The technical term for this is "not lighting money on fire because you can."

The $600 billion track nobody's using at full speed

Hyperscalers will spend over $600 billion on infrastructure in 2026, with approximately 75% targeting AI. The Stargate project alone targets $500 billion in AI infrastructure investment by 2029. Hyperscalers added $121 billion in new debt in 2025 — more than four times the average annual issuance over the previous five years. That's a lot of GPUs sitting in racks waiting for someone to justify the lease payment. The Ramp report suggests they might be waiting a while, because the companies actually using AI in production are getting better at not using all of it.

I've written before about the bill dropping 72% when Microsoft cut transcription API pricing. That was the vendor passing savings downstream. This is different. This is businesses looking at a bill that says "12 million tokens" and asking whether the thing that called the model twelve million times could have called it twice and cached the rest. The answer is usually yes. The friction is that nobody bothered to check until the bill crossed a threshold that made someone's manager ask a question.

The construction ERP doesn't call a model to rewrite every email. It calls the model once to generate a template library, caches those templates in D1, and serves them locally until a project manager edits one and decides the edit is worth saving. That decision — call once, cache many — is the difference between a $40 API bill and a $1,200 one. I made that change six months ago not because the API got more expensive, but because I finally did the arithmetic and realized I was paying for the same generation craft hundreds of times.

The hyperscalers built a $600 billion racetrack. Businesses showed up, looked at the track, and decided to drive in second gear. Not because the cars are bad — because the finish line turned out to be three miles closer than the architects thought.

The Ramp data shows the same pattern at scale. In July, the top 1% of businesses spent a median $7,400 per employee on AI, the top 10% spent $650, and the median firm spent $11.95 per employee. By September, the top spenders had reduced their per-employee burn. That's not because they stopped using AI. It's because they stopped using it wrong. The systems are still running — the ERP still rewrites estimates, the dialer still sends review requests, the agent on ciprari.ai still stages edits to my résumé. They just cost less now, because someone finally looked at the invoice and realized that calling gpt-4 to capitalize a name is not a business-critical use of a frontier model.

The mismatch here is structural. AI-related services delivered roughly $25 billion in revenue in 2025 against more than $250 billion in infrastructure spending — about 10 cents of revenue per dollar of capex. Goldman Sachs warned that maintaining returns would require $1 trillion or more in annual profit by 2026. The vendors built for a world where every business would throw frontier models at every problem. The businesses that actually deployed AI learned to throw the small model at most problems and the big one at the three that matter. That's not a failure of AI — that's AI working exactly as advertised, which is to say it works well enough that you can stop overpaying for it and it still works.

The ten platforms I've shipped all follow the same rule now: pick the smallest model that solves the problem, call it the fewest times that gets the result, and cache everything that doesn't change. The construction ERP caches document templates. The estimating tool caches material-cost explanations. The Twilio dialer caches message rewrites until a client changes their script. None of that required a rewrite — it required looking at a D1 table and asking whether the thing I just generated was identical to the thing I generated yesterday. If the answer is yes, the second call is waste. Ramp's data suggests a lot of businesses just figured that out.

The hyperscalers are still building. The models are still getting better. The bill, for people who are paying attention, is going down. That gap — between what it costs to build the infrastructure and what it costs to actually use it well — is the only number that matters in production. The rest is just architecture.


— Cole Ciprari · Business Systems Architect · Worcester, MA
my résumé is an operating system → ciprari.ai · linkedin.com/in/coleos · cole@ciprari.ai
WAS THIS ANY GOOD?
Anonymous, one tap, no account. Tap again to undo.
▚▞ GET THE NEXT RELEASE
A new release every day, plus the Sunday Rollout Report — the week's AI and tech news, summarized by a human with production access. No spam. Unsubscribe by emailing a mildly disappointed cole@ciprari.ai.
PHOSPHOR