← Back to News
ANALYSIS

Enterprise AI Spending Hits Its Efficiency Inflection as Big Customers Rein In Tokens

In late June, CNBC reported that OpenAI and Anthropic are facing a new commercial reality as their largest customers shift from tokenmaxxing to efficiency. Uber reportedly burned its annual AI budget in four months, Lindy moved all of its traffic off Claude to a cheaper model, and the FinOps community put AI tokenomics at the center of its 2026 agenda — just as both labs filed confidentially for IPOs.

By Michael Eakins•• min read
Enterprise AIAI EconomicsFinOpsOpenAIAnthropic

Executive Summary

For two years, enterprise AI spending ran on a spend-first, optimize-later logic. In the final weeks of June 2026, a cluster of reports made clear that the logic is reversing. CNBC reported on June 26 that OpenAI and Anthropic are confronting a new reality as their largest customers move from tokenmaxxing — pumping as many tokens as possible into the most capable model — toward efficiency. The supporting anecdotes were specific: Uber reportedly exhausted its entire annual AI budget in roughly four months, and the chief executive of the AI startup Lindy moved one hundred percent of the company traffic off Anthropic Claude models onto DeepSeek to cut costs. At the same time, the FinOps community reframed the AI invoice as a cost to be governed, with value per token replacing cost per token as the metric to chase.

The timing matters because both major Western labs filed confidentially for public offerings in early June. A consumption-growth narrative that flourished in private rounds now has to survive public-market scrutiny precisely as the demand side begins to discipline its token usage.

What Was Reported

Uber's annual AI budget

Spent in ~4 months

Reporting on the efficiency shift cited Uber burning through its entire annual AI budget in roughly four months — an overrun emblematic of the tokenmaxxing era now ending.

Lindy's traffic migration

100% off Claude

The CEO of AI startup Lindy moved all of the company traffic from Anthropic Claude models to DeepSeek to cut costs — a decision that was unthinkable to make loudly while the flagship was visibly better at almost everything.

Anthropic annualized run rate

~$47B in May 2026

Up from roughly ten billion dollars in revenue for all of the prior year. Analysts note current growth rates are likely the fastest these companies will ever post — partly because some of that growth came from customers who were not optimizing and now are.

The behavioral reports landed alongside a wave of cheaper model economics. Earlier in the month, Microsoft AI unveiled a suite of seven first-party models under the MAI designation, led by a reasoning model designed to match premium logical output at a far lower token cost. Budget-tier frontier-class models — DeepSeek V4 Flash among them, at roughly fifteen cents per million input tokens — now undercut the mini and nano tiers from every major Western lab while staying close enough in quality to handle the bulk of production traffic.

Why It Is Happening Now

The shift is not a sudden discovery of discipline. It is the consequence of the price of capability collapsing far enough that discipline finally pays off. When the flagship model was the only one that worked, frugality saved nothing. Now that a model a fraction of the price clears the bar for most traffic, the same frugality saves most of the bill.

The two regimes of enterprise AI spending

Tokenmaxxing eraEvery request routed to the best model on the menu, unmetered developer keys, cost reconciled quarterly if at all. Output quality was the only tracked metric.
Efficiency eraA cheap router downshifts routine traffic to a budget tier and escalates only the hard tail; caching discounts repeat traffic; FinOps governs the total bill, not just the token line.
What changed between themCapable models got cheap, AI budgets moved from innovation funds to scrutinized operating costs, and the invoices grew large enough to attract people whose job is to ask whether spend produced value.

The FinOps framing adds a structural point: the token invoice is only one of several cost buckets. Self-managed inference compute, data and retrieval costs, human review, and engineering maintenance all sit alongside it. A team that drives its token bill to zero by self-hosting and then spends three engineers babysitting a cluster has not saved money — it has moved the cost somewhere nobody is counting.

The Stakes For The Labs

The efficiency inflection, in sequence

Jun 2, 2026

Microsoft ships in-house MAI models

Seven first-party models led by a reasoning model designed to match premium output at far lower cost — a bet that buyers want cheaper capability, not just more of it.

Early Jun 2026

OpenAI and Anthropic file confidentially for IPOs

Both labs move toward public markets just as the demand-side repricing begins, putting their consumption-driven growth curves under harsher scrutiny.

Jun 26, 2026

CNBC reports the tokenmaxxing-to-efficiency shift

Coverage names the trend, with Uber and Lindy as concrete examples and analysts warning current growth rates are the fastest the labs will ever post.

Jun 2026

FinOps puts AI tokenomics at the center

The community reframes the AI invoice as one of nine cost buckets to be governed, with value per token as the metric to chase.

The labs are not in danger; they are growing at rates most companies never see. The question the efficiency turn raises is whether the steepest part of the curve is already behind them. When the biggest accounts move from unmetered keys to governed budgets, the labs do not lose them — but they lose the portion of growth that came from waste. The defensive move is already visible: ship your own cheap tier so the downshift happens inside your product rather than to a competitor. It retains the logo while shrinking the invoice, which is why every major provider now has an aggressively priced budget SKU.

For the full strategic analysis of how the demand side is repricing the frontier — the unit economics, the builder playbook, and the Jevons counterargument — see the companion feature, The Efficiency Turn.

What To Watch

Three signals will indicate whether this is a durable repricing or a quarter of belt-tightening. First, deep and proactive cuts to flagship API pricing — not just cheaper new tiers — would confirm the demand side now sets the price. Second, a widening gap between user growth and revenue growth in any eventual IPO disclosure would show the efficiency turn reaching the financials. Third, whether budget-tier models keep closing the quality gap, since the entire shift rests on cheap models being good enough. A sudden re-widening at the frontier would hand pricing power back to the flagships.

Sources

  • CNBC, "OpenAI and Anthropic face new AI reality as users shift from tokenmaxxing to efficiency" (June 26, 2026)
  • CNBC, "Microsoft unveils new AI models to lessen reliance on OpenAI and lower costs for developers" (June 2, 2026)
  • FinOps Foundation, FinOps X 2026 coverage on AI token economics
  • Public reporting on DeepSeek V4 Flash API pricing and budget-tier model comparisons (2026)