AI’s new paradox: Cheaper models, bigger bills

Published on the 27/08/2026 | Written by Heather Wright


The Inference Paradox and how to cut those bills…

AI models are getting cheaper, but AI bills are heading in the opposite direction.

That’s the contradiction at the heart of new Gartner research, which predicts AI inference costs per agentic workflow will increase more than fivefold by 2028 as organisations move from simple chatbots to more sophisticated AI agents capable of reasoning, planning and executing multistep tasks. The analyst company calls it the ‘inference paradox’ with improving model economics being overwhelmed by increasingly complex AI workloads.

While model prices might be falling, that’s luring users to build ever more complex workflows – and the greater token consumption of those workflows can outweigh savings from reducing model prices and escalating inference costs.

“The harsh economics of the Inference Paradox are exemplified by the differences between a simple chatbot and an AI agent,” Gartner senior director analyst Will Sommer says. “Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate and question itself.”

All of that processing comes at a cost.

Gartner says routing a task to an agentic reasoning model increases inference costs by at least five times compared with a basic chatbot interaction, and often much more as complexity grows. At the same time, organisations are discovering that increasingly capable AI systems consume vastly more tokens than traditional conversational AI.

Gartner argues many organisations are making matters worse by allowing inefficient agent behaviour to creep into systems as they scale.

“There is an enormous pool of waste in any agentic system,” Gartner says in a research note. “Most of the bill in agentic systems is redundancy.”

Among the issues identified in 50% of Token Costs Can Be Cut for AI Agents With Zero Quality Loss are repeated queries, redundant processing, oversized context windows and agent loops that repeatedly consume tokens without adding value. Gartner estimates roughly 31 percent of production queries repeat work organisations have already paid for. Agents consume around 100 input tokens for every output token generated, while 40 percent to 60 percent of agent tool output tokens can often be removed without any loss of quality.

This is where Gartner believes the next AI battleground is emerging as companies move beyond the first phase of experimentation and the second of governance, trust and responsible deployment and into AI engineering efficiency, or building systems that deliver the same outcomes using fewer resources.

That starts with model selection.

“Most tasks do not require frontier intelligence,” the report notes, warning that defaulting every workload to the largest and most capable models is ‘immensely wasteful’. Instead, Gartner recommends routing requests to the cheapest model capable of completing a particular task and escalating only when more sophisticated reasoning is genuinely required.

The report report estimates organisations can reduce costs by up to 60 percent through routing and workload optimisation alone.

Gartner argues that visibility becomes increasingly important as organisations scale AI into production. It recommends tracking metrics including cost per completed task, token consumption by model, spend by customer segment and the relationship between AI expenditure and customer outcomes. Without detailed measurement, organisations often fail to identify expensive models performing basic work or AI features generating little business value.

At the centre of Gartner’s recommendations is a concept that will sound familiar to cloud veterans: AI FinOps.

The company recommends introducing AI gateways, workload routing, caching, attribution and budget controls to actively manage AI consumption rather than simply paying whatever bill arrives at the end of the month. It describes routing, caching and orchestration as critical to preventing costs from spiralling as agent complexity increases.

“Hold agents to high-value tasks and delegate the rest down a tier,” the report says.

“Product leaders cannot rely on more efficient token economics to rationalise AI costs,” Sommer says. “Each successive generation of AI capability will necessitate more, and often more expensive, tokens.” He says there is ‘no reliable, economical one-size-fits-all model on the horizon’ and that organisations will increasingly need to manage complex multimodel environments.

The move by some AI providers to usage-based billing – as previously reported by iStart – has exacerbated issues.

Gartner’s report also identifies opportunities inside the architecture of agentic systems themselves.

Rather than feeding large volumes of information into models, the report recommends retrieval-based approaches that selectively surface only relevant content. Using retrieval-augmented generation (Rag), for example, can reduce token consumption by up to 75 percent while maintaining accuracy. Context compression approaches can reduce token volumes by 60 percent to 95 percent, while redesigning agent loops can remove 40 percent to 60 percent of redundant processing.

The broader message is that AI economics are changing.

For the past two years, the primary question has been whether AI could perform useful work. Gartner suggests the more important question now may be whether organisations can afford to run increasingly sophisticated AI systems at scale – and how to control what that intelligence costs.

Post a comment or question...

Your email address will not be published.

This site uses Akismet to reduce spam. Learn how your comment data is processed.

MORE NEWS:

Processing...
Thank you! Your subscription has been confirmed. You'll hear from us soon.
Follow iStart to keep up to date with the latest news and views...
ErrorHere