Tokenomics: Why making AI pay is tricky

If you have used a free version of an ChatGPT or its AI rivals, then you are obviously getting a good deal.
Firms like Microsoft, Google and Anthropic have invested hundreds of billions of dollars in developing Large Language Models (LLMs) the tech behind those services.
So getting, ChatGPT, Claude or Gemini to help with your speech or holiday plans is a bargain.
But, naturally, those firms want to recoup their investment, so they offer paid-for versions of their AI, which have extra features for tasks like coding or billing.
Meanwhile, third party firms are building and selling services based on AI agents, usually based on an LLM, which are trained to do specific tasks.
But setting a price for those services is surprisingly difficult.
"Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know," says Simon Gooch at Saviynt, an identity management company which is incorporating agentic AI into its services.
That's because of rapidly changing economics around tokens, the building blocks of LLMs and agentic AI.
When a user asks an LLM, like ChatGPT or Anthropic's Claude to answer a question, generate software code, or automate a process, that prompt is broken down into mathematical chunks called tokens, which can be processed by the model.
The LLM's response also comes in the form of tokens, which are converted back into text, software code, or a set of commands to automate a process.
The problem is this process is not entirely predictable.
Subtle variations in the prompt can produce different answers. The same prompt will not always produce the same answer. Different models will produce different answers.
Meanwhile, in agentic systems, businesses use multiple AI agents together to make decisions and take actions, further increasing both token use and unpredictability.
While the cost of individual tokens – or the credits used to pay for them - has plummeted in recent years, according to analysis by Goldman Sachs, the number of tokens consumed by businesses, and consumers, has skyrocketed.
The bank forecasts that, external token consumption will increase 24 times between 2026 and 2030 to 120 quadrillion tokens a month, as companies shift from to use AI agents.
But companies, and individuals, using AI systems often have a tenuous grasp on just how many tokens they are burning through – until they either run out or get their monthly bill.
Even Microsoft has reportedly reined back, external its engineers' use of some third party coding tools, while Uber apparently tore through, external its AI coding token budget for a year in a matter of months earlier this year.
Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, said companies can be caught out as they experiment with or implement AI internally, as staff burn through tokens.
"People are finding it really hard to manage that cost… it's a non-deterministic output, so it's a non-deterministic value," he said.
Companies are finding ways to work around this.
Oliver King-Smith, founder of engineering software firm smartR AI, says smaller organizations can "can fly under the radar and use [flat fee] personal accounts which I am sure the big vendors don't like."
But, he says, "This has to end at some point in time, because the big guys are taking a bath on those accounts."
Once the big AI platforms start facing pressure from shareholders to show a profit, he predicts: "They will start clamping down."
King-Smith says companies should also think more carefully about what AI models to use.
Companies also needed to be much more precise with their prompts, says Rob Steele, CFO at UK accounting software firm iplicit.
"You wouldn't send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?"
The situation can become difficult to control when companies build AI into a product that could be rolled out to thousands of users, Venters points out.
AI costs could start to balloon. For example, managers may realise they need tokens not just for core software development, but for other tasks such as testing, security, or for implementing guard rails.
"It's particularly hard when you're looking at agentic processes," Ventners says.
Employing more AI agents can be done with the click of a button, whereas expanding the human workforce would involve careful discussions over headcount and hiring, he says.
Venters points out, while token costs might be unpredictable, it might be that the company is ultimately getting more value from their token use with AI.
"It's not quite the same as a calculator," he says. "The more you give it, the more expensive it is, but the better the result may be."
But companies still need to pass those costs onto their own customers.
"Nobody's really figured it out," says Bill Peterson, senior director of product marketing, at Sumo Logic.
The software firm is previewing new security services based on agentic AI, he explains, but is in discussion with corporate customers about how to charge for them.
"We're still having some fun conversations about this internally," he says drily.
Options could include simply raising prices across the board, he says, paying by results, or charging for "bundles" of incidents.
But whatever price structure it chooses could be upended if and when the large language model providers change their own pricing strategies.
"You get into variable pricing, and it's changing every couple of months" he says. "Customers don't like that. That's not how anybody builds a budget."
Read the full story at BBC ↗
AI providers like Microsoft, Google and Anthropic have invested heavily in large language models and offer paid versions alongside free tiers. Third-party companies building AI services on top of these models face a novel problem: pricing. Token consumption—the computational cost of processing AI prompts and responses—is difficult to predict because identical prompts can yield different results, and multi-agent systems amplify this variability. While individual token costs have fallen, overall consumption is rising steeply; Goldman Sachs projects monthly token use will increase 24-fold by 2030. Companies and individuals often lack visibility into their actual token spend until billing arrives, leading to surprises like Microsoft reining in coding tool use and Uber exhausting annual budgets in months. Experts note that non-deterministic output creates non-deterministic cost, complicating budget forecasting. Some smaller organisations use personal flat-fee accounts to sidestep this, but industry observers expect vendors to close this gap once profit pressure intensifies. Service providers considering AI integration face difficult choices: raising prices broadly, charging per result, or bundling charges—all vulnerable to shifts in underlying token pricing. No consensus pricing model has emerged.
Read the full story at BBC ↗
If you have used a free version of an ChatGPT or its AI rivals, then you are obviously getting a good deal.
Firms like Microsoft, Google and Anthropic have invested hundreds of billions of dollars in developing Large Language Models (LLMs) the tech behind those services.
So getting, ChatGPT, Claude or Gemini to help with your speech or holiday plans is a bargain.
But, naturally, those firms want to recoup their investment, so they offer paid-for versions of their AI, which have extra features for tasks like coding or billing.
Meanwhile, third party firms are building and selling services based on AI agents, usually based on an LLM, which are trained to do specific tasks.
But setting a price for those services is surprisingly difficult.
"Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know," says Simon Gooch at Saviynt, an identity management company which is incorporating agentic AI into its services.
That's because of rapidly changing economics around tokens, the building blocks of LLMs and agentic AI.
When a user asks an LLM, like ChatGPT or Anthropic's Claude to answer a question, generate software code, or automate a process, that prompt is broken down into mathematical chunks called tokens, which can be processed by the model.
The LLM's response also comes in the form of tokens, which are converted back into text, software code, or a set of commands to automate a process.
The problem is this process is not entirely predictable.
Subtle variations in the prompt can produce different answers. The same prompt will not always produce the same answer. Different models will produce different answers.
Meanwhile, in agentic systems, businesses use multiple AI agents together to make decisions and take actions, further increasing both token use and unpredictability.
While the cost of individual tokens – or the credits used to pay for them - has plummeted in recent years, according to analysis by Goldman Sachs, the number of tokens consumed by businesses, and consumers, has skyrocketed.
The bank forecasts that, external token consumption will increase 24 times between 2026 and 2030 to 120 quadrillion tokens a month, as companies shift from to use AI agents.
But companies, and individuals, using AI systems often have a tenuous grasp on just how many tokens they are burning through – until they either run out or get their monthly bill.
Even Microsoft has reportedly reined back, external its engineers' use of some third party coding tools, while Uber apparently tore through, external its AI coding token budget for a year in a matter of months earlier this year.
Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, said companies can be caught out as they experiment with or implement AI internally, as staff burn through tokens.
"People are finding it really hard to manage that cost… it's a non-deterministic output, so it's a non-deterministic value," he said.
Companies are finding ways to work around this.
Oliver King-Smith, founder of engineering software firm smartR AI, says smaller organizations can "can fly under the radar and use [flat fee] personal accounts which I am sure the big vendors don't like."
But, he says, "This has to end at some point in time, because the big guys are taking a bath on those accounts."
Once the big AI platforms start facing pressure from shareholders to show a profit, he predicts: "They will start clamping down."
King-Smith says companies should also think more carefully about what AI models to use.
Companies also needed to be much more precise with their prompts, says Rob Steele, CFO at UK accounting software firm iplicit.
"You wouldn't send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?"
The situation can become difficult to control when companies build AI into a product that could be rolled out to thousands of users, Venters points out.
AI costs could start to balloon. For example, managers may realise they need tokens not just for core software development, but for other tasks such as testing, security, or for implementing guard rails.
"It's particularly hard when you're looking at agentic processes," Ventners says.
Employing more AI agents can be done with the click of a button, whereas expanding the human workforce would involve careful discussions over headcount and hiring, he says.
Venters points out, while token costs might be unpredictable, it might be that the company is ultimately getting more value from their token use with AI.
"It's not quite the same as a calculator," he says. "The more you give it, the more expensive it is, but the better the result may be."
But companies still need to pass those costs onto their own customers.
"Nobody's really figured it out," says Bill Peterson, senior director of product marketing, at Sumo Logic.
The software firm is previewing new security services based on agentic AI, he explains, but is in discussion with corporate customers about how to charge for them.
"We're still having some fun conversations about this internally," he says drily.
Options could include simply raising prices across the board, he says, paying by results, or charging for "bundles" of incidents.
But whatever price structure it chooses could be upended if and when the large language model providers change their own pricing strategies.
"You get into variable pricing, and it's changing every couple of months" he says. "Customers don't like that. That's not how anybody builds a budget."
Read the full story at BBC ↗
Microsoft, Google and Anthropic have invested hundreds of billions of dollars in developing large language models. Free AI services like ChatGPT, Claude and Gemini represent good value for users given the infrastructure investment. Tokens are the mathematical building blocks that break down prompts and responses in large language models and agentic AI systems. The same prompt does not always produce the same answer from an LLM, and different models produce different answers. Multi-agent agentic systems increase both token use and unpredictability compared to single-model systems. Individual token costs have fallen in recent years while overall token consumption by businesses and consumers has skyrocketed. Goldman Sachs forecasts token consumption will increase 24 times between 2026 and 2030 to 120 quadrillion tokens per month. Microsoft has reined back engineers' use of third-party coding tools due to token cost concerns. Uber exhausted its annual AI coding token budget in a matter of months earlier this year. Companies lack clear visibility into token spending until they run out or receive bills. The non-deterministic nature of AI outputs creates non-deterministic value, making cost management difficult. Smaller organisations currently use flat-fee personal accounts to bypass usage-based pricing models. Large AI platform vendors will clamp down on flat-fee workarounds once facing shareholder pressure to show profit. More precise prompting and careful model selection can help reduce token consumption. Building AI into products for thousands of users makes token costs difficult to control and can lead to unexpected cost expansion. Expanding AI agent deployment can be done with a click, whereas expanding human workforce requires structured discussions. Increased token use may yield better results, making higher consumption potentially justified by improved value. Service providers are exploring pricing options including broad price increases, per-result charging, and bundled incident charging. No consensus pricing model for AI services to end customers has emerged across the industry. Variable pricing from LLM providers changing every couple of months frustrates customers' budgeting efforts.
Read the full story at BBC ↗
- AI services require pricing models, but token consumption is unpredictable—same prompts produce different outputs, and multi-agent systems compound this variability
- Token costs have fallen while consumption has risen sharply; Goldman Sachs forecasts 24× increase in monthly token use by 2030 as businesses scale AI agents
- Companies lack visibility into token spending until bills arrive; Microsoft and Uber have both hit unexpected costs, creating budgeting challenges
- Vendors currently tolerate flat-fee workarounds but may tighten pricing as shareholder pressure for profitability increases
- No industry standard has emerged for pricing AI services to end customers—options range from bulk price rises to per-result or bundled incident charging