Jump to content
  • Blog, Artificial Intelligence
  • Published on: 27.08.2026
  • 18min

From Token Panic to Cost Control

Tokenmaxxing Was Never a Strategy

Another word of warning before we start: this is long. Longer than the attention economy recommends. If you came for a bite-sized LinkedIn take, this may not be for you. What follows is the uncomfortable version: field notes on what AI actually costs, minus the hype and minus the flattery. Written for people who would rather know the number than feel good about it. Ears up. 

In July 2026, Sam Altman sat on a cream-colored couch in San Francisco, in front of a room full of enterprise customers, and called the cost of AI "almost a meme". His paraphrase of what those customers now tell him: "My company already used up its entire IT budget for 2026 in the first quarter. Can you make this more efficient?"(1) The man who lit the fire three and a half years ago was, politely, admitting the room had gotten too hot.

Across MHP’s work with enterprise clients, the problem usually becomes visible once the invoice has stopped being a rounding error and started being a board topic. The pattern is by now familiar. The spend is real, it is growing, and almost nobody can see it while it happens. Someone eventually asks how it got so high. Nobody can answer, because the meter ran the whole time and no one was watching.

The scale explains the nerves. Companies worldwide now spend more than a trillion dollars a year on AI (2). A KPMG survey of about 2,100 organizations found that only 35 percent can fully track and monitor that spend, while 13 percent find out the number when the invoice lands, which is the one moment when it is already too late to change anything (3). Citrini Research has a name for the mood in the boardroom: token panic (4). A year ago, the same rooms were celebrating high token consumption as proof of ambition. That was tokenmaxxing, and it was never a strategy. It was a phase, and the phase is ending.

The thesis of this piece is straightforward:

  • AI FinOps - the discipline of governing what AI actually costs and what it returns - is not yet on most companies' radar as a discipline at all. That will change.

From 2027 onwards, strong AI FinOps capabilities will become a real competitive differentiator. Companies that can answer "what did that token bill buy us" will out-invest the ones that cannot, because they will know where to spend and where to stop.

For CIOs, CTOs, CFOs and AI platform leaders, the issue now reaches far beyond tooling. Governing AI economics has become a management responsibility.

"The conversation is evolving. First, organizations wanted to scale AI. Now they need to scale it responsibly, economically, and with clear business outcomes. Companies that master this balance will turn AI from a technology investment into a sustainable competitive advantage." - Björn Kasten, Partner, MHP

AI FinOps reaches beyond cost control. Digital sovereignty also depends on knowing what AI consumption costs, which value it creates and where dependencies arise. Organizations that have this transparency can make model, provider and architecture decisions from a position of control.

Three problems, one root

When we look at how companies got here, the same three problems show up almost every time.

The first is a governance problem: AI tooling gets waved through as an experiment, and the FinOps structures that should come with it, budgets, ownership, metrics, never get thought through. The costs start running immediately. The oversight arrives months later, if at all. An experiment with no cost owner is not an experiment. It is an open tab.

The second is an awareness problem: AI FinOps is barely established as a discipline. Most organizations that would never dream of running cloud without cost governance are running AI exactly that way, because they have not yet registered that it is the same problem in a faster costume. This gap is likely to close rapidly around 2027, and the companies that close it early will have a genuine head start.

The third is a value-measurement problem: and it is the most consequential one. Costs are becoming visible through licenses and token consumption, all the way to the monthly bill. The value stays a black box. Companies can see what they spend and still cannot say whether the investment carries itself. Seeing the cost is progress. It is also only half the equation, and the cheaper half.

The root underneath all three is the same. Companies treated a metered technology like a flat-rate one. Everything that follows is the bill for that assumption.

Managers now face a clear choice. They can continue treating AI consumption as invisible overhead, preserving short-term convenience. Or they can manage it as a strategic resource and create the transparency needed to scale AI with control.

This is a cloud story wearing a new hat

  • The part that gets lost in the noise: most AI spend is not an exotic new budget line. It is Cloud!

Gartner expects AI infrastructure alone to grow from 965 billion dollars in 2025 to almost 1.75 trillion by 2027, by far the largest slice of total AI spend (2). Every token a model processes is compute running in someone's data center, metered by the second.Storage, networking, GPU capacity, inference endpoints: the bill arrives through the same cloud accounts finance was already trying to keep honest.

Which is why the discipline the industry spent a decade building for cloud is the discipline AI needs now. FinOps taught us to tag spend, allocate it to the team that caused it, right-size the resources, set budgets, and put the cost in front of the person making the decision. All of it transfers. Instances become models. Instance hours become tokens. Reserved capacity becomes provisioned throughput.

  • The core equation has not moved: cost is price times quantity, and it comes down by reducing one or the other.

Eine Organisation mit ausgereiftem Cloud FinOps verfügt bereits über eine Grundlage. Sie kann eine etablierte Disziplin auf eine schnellere und weniger verzeihende Kostenbasis anwenden.

Dennoch durchbricht KI viele der komfortablen Gewohnheiten dieser Disziplin. Cloud-Preise waren vergleichsweise stabil und veröffentlicht. KI-Dienste verwenden unterschiedliche Abrechnungsmodelle, Modellportfolios verändern sich schnell, und tokenbasierte Kosten lassen sich mit traditionellen Finanzwerkzeugen nur schwer zuordnen. Finanzteams steuern daher Kosten, deren Mechanismen oft tief in technischen Systemen und der Dokumentation der Anbieter verborgen liegen.

If an organisation already runs cloud FinOps well, they are not starting from zero. They are pointing a muscle they already have at a faster, less forgiving target. That is the genuinely encouraging part.

The discouraging part is that AI breaks the comfortable habits of the discipline. Cloud prices were relatively stable and published. AI pricing moves, sometimes overnight, and not always downward. Anthropic's Fable 5 priced at ten dollars per million input tokens and fifty dollars per million output tokens, twice the standard token price of Claude opus 4.8 (5). And that is the moderate end of the frontier. OpenAI's GPT-5.5 Pro sits at thirty dollars per million input and a hundred and eighty for output. That is three to nearly four times as much, which makes Fable 5's record-setting rates look almost restrained (6).

The sharpest increases hide in the mechanics rather than the price list: Anthropic's newer models split the same text into roughly 30 percent more tokens than before. Same task, more billable units, an increase that lives in the developer documentation where most finance teams will never look.(7)

From visibility to guardrails

The trigger for most of the panic was a shift in how AI gets billed. Fixed-price subscriptions are giving way to consumption pricing, where every request costs something and every request costs a different amount depending on the model, the context it drags along, and how many steps an agent takes before it stops. With autonomous agents, a single task fans out into dozens of billable calls. The cost stops being a number anyone can estimate and becomes a number they can only read after the fact.

Across MHP’s AI FinOps engagements, the first move is consistently the same: creating visibility before adding another tool.

We did not start by buying anything. We switched on the cost signals that already existed and nobody was looking at: the live cost meter, the per-session readout, the exact-cost hover in the editor. When a developer could watch a single session tick past a real number in real time, behavior changed without a single policy memo. Cost stopped being an abstraction that lands on someone else's desk and became something people felt while they worked. That matters, because only 53 percent of companies run cost dashboards for AI and only 40 percent set usage budgets (3). Closing that gap was the cheapest win on the table, and we took it first.

The second movewas right-sizing, which came down to one blunt question: does this task actually need the most expensive model? We had seen how quickly grabbing the biggest model by default could become one of the most expensive habits in an AI project. So, we introduced routing. Simple requests such as summaries, classification and boilerplate went to smaller or open-source models. Hard reasoning and complex code went to the top tier, on purpose rather than by reflex. The lesson was not that AI was too expensive. Using a Formula 1 engine to drive to the bakery was always going to hurt.

The third move was guardrails. A shared budget invites a tragedy of the commons, so we capped per-user consumption with a clear path to raise the limit when someone genuinely needed it. That is a circuit breaker, not a punishment. We set stop conditions, so an agent ends when the acceptance criterion is met, because an agent left open-ended keeps working, and keeps billing, long past the point anyone needed. We trimmed context, because bloated context is money set on fire before the model even starts. None of it was glamorous. All of it worked.

The most relevant outcome is behavioral. People start making different decisions. The model picker stopped being decoration. Cost moved to the left, into the moment of choosing how to solve a problem, instead of arriving as a shock at month end. And the conversation shifted from "how do we stop people using AI" to "how do we get more out of every token", which is the only version of that conversation worth having.

That shift in behavior eventually changes the architecture itself. Once teams stop choosing models by reflex, routing becomes more than a tactical cost lever. It reflects a larger principle: model choice is becoming an architecture decision with lasting purchasing consequences. Tasks inside an agent system differ in the reasoning power, speed and cost they require. In industrial environments, these choices compound across engineering, production, procurement, logistics and aftersales. A decision made once in the architecture can shape the cost and dependency profile of thousands of recurring processes.

The measurement problem, on three levels

Visibility, routing and guardrails buy cost control. They do not buy the thing everyone actually wants, which is knowing whether the money bought anything. This is where the discipline runs out of road, and the road runs out on three levels.

First, the effort is hidden. The prompt work, the iteration until the output is usable, the quality check afterward: none of it appears on a license invoice, and all of it eats real capacity. A tool can look cheap on the bill while quietly consuming a developer's afternoon. That afternoon is a cost too. It just does not have a line item.

Second, efficiency is claimed, not proven. Vendors promise "10x productivity". Almost nobody measures whether the generated code was actually production-ready, or whether the review effort to make it safe canceled out the gain. A number that impressive should come with a method. It rarely does.

Third, quality is not a constant. Fable 5, GLM-5.2 or GPT-5.5 might ship better code today than in three months, or worse, depending on the next update. Companies have no stable baseline to measure against, so even an honest before-and-after comparison sits on shifting ground.

Put the three together and the core problem is uncomfortable. The industry is selling efficiency it cannot measure, and the buyers are paying for it without knowing what they get. That is not a sustainable trade for long, which is exactly why the ability to measure value is about to become worth more than the ability to cut cost.

Provider switching is only one part of sovereignty. Real sovereignty appears when a company understands the economic and operational consequences of its choices well enough to act.

What AI cost control requires now

The cloud taught this discipline once. Tag it, right-size it, budget it, put the cost where the decision gets made. AI is asking whether anyone was paying attention. An organization that built real FinOps habits for cloud already has most of what it needs, and the metering just got faster and less forgiving.

The gaps in governance, awareness and measurement are all closing, and they will not close at the same speed for everyone. That creates an opportunity.

From 2027, AI FinOps stops being back-office hygiene and will increasingly separate disciplined companies that manage AI as a strategic resource from those still paying to feel busy. Frontier models are worth every credit for the hard problems, and most daily work does not live at the frontier. Knowing the difference, and building the visibility, the architecture and the guardrails to act on it, is what separates a team that masters AI cost from one that just pays it.

Ein Produkt, das ein Gefühl von Dynamik erzeugt, während Kosten und Wert undurchsichtig bleiben, ist kein gutes Geschäft. „Tokenmaxxing“ war eine Eitelkeitskennzahl, die sich als Dynamik verkleidete. Was sie ersetzt, ist weniger aufregend und weitaus wertvoller: ein Unternehmen, das weiß, was es ausgibt, warum es das ausgibt und was es dafür zurückbekommt.

Palantir's chief executive Alex Karp was provoking when he said a product that makes you feel smart while your company goes broke is not a good deal (8). He was also not wrong. Tokenmaxxing was a vanity metric dressed up as momentum. What replaces it is less exciting and far more valuable: a company knowing what it spends, why it spends it, and what it got back.

Quellen

  1. OpenAI. „Intelligence at Work.“ OpenAI Livestream, July 2026.  https://openai.com/business/intelligence-at-work/
  2. Gartner. „Gartner Says Worldwide AI Spending Will Total $2.5 Trillion in 2026“, January 15 2026. www.gartner.com/en/newsroom/press-releases/2026-1-15-gartner-says-worldwide-ai-spending-will-total-2-point-5-trillion-dollars-in-2026
  3. KPMG International. Global AI Pulse Q2 2026, June 2026. assets.kpmg.com/content/dam/kpmgsites/xx/pdf/2026/06/global-ai-pulse-q2.pdf
  4. Citrini Research. „State of the Themes: June 2026“, 8. Juni 2026, section „Token Panic“. www.citriniresearch.com/p/state-of-the-themes-june-2026
  5. Anthropic. „Claude Fable 5: Availability and Pricing.“ Anthropic, Accessed July 22, 2026. www.anthropic.com/claude/fable
  6. OpenAI. “Introducing GPT-5.5.” April 23, 2026. See also: OpenAI, “GPT-5.5 Pro Model,” OpenAI API Documentation. Accessed July 22, 2026. openai.com/index/introducing-gpt-5-5/
  7. Anthropic. “Token Counting.” Claude Platform Documentation. Accessed July 22, 2026. platform.claude.com/docs/en/build-with-claude/token-counting
  8. CNBC. “Palantir CEO Alex Karp Slams OpenAI and Anthropic Over Token-Based AI Pricing.” CNBC, July 1, 2026. Accessed July 22, 2026. https://www.cnbc.com/2026/07/01/palantir-karp-open-ai-anthropic-tokens.html

Your contact

Speak with our experts today. We look forward to hearing from you.

bjoern.kasten(at)mhp.com

LinkedIn

Björn Kasten

Digital Core & Technology | MHP