Skip to main content

I Is Getting Cheaper. So Why Are AI Costs Becoming a Bigger Problem?

AI Operating Costs and Tokenomics Shift 2026

AI Is Getting Cheaper. So Why Are AI Costs Becoming a Bigger Problem?

AI is getting cheaper.

Model prices are falling. Open-source alternatives are multiplying. Smaller models are getting better. Inference is becoming more efficient.

That should make AI less expensive to run.

And yet McKinsey’s 2026 State of AI survey finds that one in five organizations is already limiting AI use because of operating costs.

At the same time, 60% expect to increase AI investment next year.

So we have three things happening at once:

  • AI is getting cheaper.
  • Companies are spending more.
  • Cost is becoming a bigger constraint.

That sounds contradictory.

It is not.

The economics are simply moving from a price problem to a management problem.

Cheaper tokens do not mean cheaper AI

McKinsey senior fellow Michael Chui points to an emerging conversation among CFOs and AI leaders around “tokenomics.”

The logic is simple:

Price per unit ↓
while:
Usage ↑↑
which can still produce:
Total cost ↑

As AI gets cheaper, organizations use more of it.

More reasoning. More agents. More coding. More workflows. More employees. More retries. More inference.

The cheaper intelligence becomes, the easier it is to consume a lot of it.

Cloud computing has taught us this lesson before. A falling unit price is not much comfort if consumption grows faster.

Scale changes everything

At pilot scale, cost discipline is easy to ignore.

A team uses a frontier model because it works. A prototype runs a few thousand calls. An agent retries more often than expected. Nobody loses sleep over it.

Then the system works. And the company scales it.

Now multiply that across:

  • hundreds of employees
  • dozens of workflows
  • several models
  • autonomous agents
  • retrieval systems
  • cloud infrastructure
  • observability
  • human review

Suddenly AI is no longer a neat experiment. It is a recurring operating expense.

McKinsey reports that nearly nine in ten respondents use AI in at least one business function, with 44% reaching the scaling phase. Among enterprises with more than $1 billion in revenue, 40% report scaling AI agents.

The economics become much harder to ignore once the experiment succeeds.

Cost pressure may be a sign of maturity

This is where the story gets interesting.

AI high performers are not immune to cost pressure. In some cases, they encounter it sooner.

McKinsey reports that operating costs constrain high performers’ use of software coding agents at roughly three times the rate reported by other organizations.

That does not necessarily mean the technology is failing. It may mean they are using enough AI for the economics to become real.

The organizations that scale first are often the first to discover that “AI is cheap” and “AI is inexpensive to operate at enterprise scale” are not the same statement.

The real question is not which model is cheapest

Once AI becomes a portfolio of operating costs, token price is only part of the picture.

A mature cost discussion has to include:

  • inference volume
  • context size
  • retrieval
  • cloud infrastructure
  • storage
  • observability
  • human review
  • licensing
  • engineering
  • agent retries
  • governance

At that point, asking: Which model costs less? is too narrow.

The better question is: Which architecture delivers enough value to justify its total cost?

That leads to a principle I think will become increasingly important:

Do not pay for more intelligence than the decision requires.

Cheapest sufficient intelligence

Not every problem needs a frontier reasoning model. Some do. Many do not.

A high-stakes strategic analysis may justify the best model available. Routine summarization may work perfectly well with a smaller model. A classification problem may not need an LLM at all. A structured business rule may be better handled by deterministic software.

So instead of routing every task to the smartest model, organizations can think in terms of:

Business requirement

Capability needed

Risk / confidence threshold

Cost tolerance

Model or workflow choice

That is cost governance. Not “use the cheapest model.” Use the cheapest sufficient architecture.

AI is also changing the build-versus-buy equation

McKinsey found something else worth watching.

Nearly one-third of respondents say their organizations decided against purchasing at least one software product because they could build the functionality internally using AI coding tools. In technology, the figure reaches 41%.

That is not just an AI productivity story. It changes the economics of software itself.

The old question: Should we buy this tool? may increasingly become: Should we buy it, build it, customize something smaller, or not build it at all?

That decision affects:

  • licensing
  • engineering
  • maintenance
  • technical debt
  • ownership
  • integration

AI is not merely adding another technology expense. It may be reshuffling the technology budget.

Cost only makes sense next to value

A $100,000 AI system is not automatically cheaper than a $500,000 one.

If the first creates $50,000 of value, it is expensive. If the second creates $3 million, it may deserve more capital.

That is why cost governance has to connect operating expense to business outcome.

A useful portfolio view might look like:

AI Operating Cost

Realized Value

ROI

Trend + Risk

Decision

Scale | Route Cheaper | Redesign | Pause | Stop

Cost is not the decision. It is one input into the decision.

This is where Decision Systems fit

At Decision Systems AI, this is one of the problems I have been building around.

Instead of optimizing cost after an architecture is already in production, a Decision System can ask earlier:

  • Does this step need AI?
  • Does it need this model?
  • Can deterministic logic handle part of it?
  • Should the agent run continuously?
  • Is the workflow creating enough value to justify operating it at scale?

Those are architectural questions, not billing questions. And that distinction matters.

The cheapest token is not much help if the workflow should never have been designed that way in the first place.

The executive question is changing

AI will probably keep getting cheaper at the unit level. Organizations will probably keep using more of it. Both can be true.

That means declining model prices will not eliminate cost governance. They may make it more important.

Because when intelligence becomes easier to consume, the organization needs a better system for deciding where consuming more intelligence is actually worth it.

So the question is no longer: How cheap can we make AI?

It is: Are we using the right amount of intelligence, in the right workflow, at the right cost, for the value the decision actually requires?