The price of AI is falling; why are enterprises still spending more? Explained
Account subscription benefits alongside Premium Stories, Editorials, Opinions and more. Unlock these with Subscription
Robot playing the piano and composing music, generative AI for music concept | Photo Credit: ELENABS
Artificial intelligence (AI) is turning sharply cheaper to run on a per-token basis. Yet, the amount companies spend on AI is rising as they move from experiments and chatbots to more complex, multi-step applications.
During a recent executive briefing, Coforge management said that over the previous seven months, the unit cost of tokens needed to deliver the same level of AI intelligence had fallen about 100-fold, while token consumption had grown about 8,000-fold.
The company also said open-weight models accounted for about 35% of token consumption on the platforms it was observing, compared with about 11–12% in early 2025.
There are broadly two types of models. Closed or proprietary models such as OpenAI’s GPT, Anthropic’s Claude and Google’s Gemini, keep their trained parameters under the developer’s control and are typically accessed through a hosted service or API.
Open-weight models, such as Meta’s Llama, Google’s Gemma and Mistral’s open-weight models, make their trained parameters available for organisations to download and run themselves, subject to the model’s licence. Open-weight does not necessarily mean fully open-source. Their training data, training process and other components may remain undisclosed.
Coforge also argued that the number of tokens used is not, by itself, the right way to assess the economics of AI. The company said the more meaningful measure is the “cost per unit task”. This means what it costs to complete a defined piece of work and what value that work creates.
Stanford’s 2025 AI Index, an annual report tracking major developments and trends in artificial intelligence, found that the cost of querying a model with GPT-3.5-level performance on a standard AI knowledge and reasoning benchmark fell from $20 per million tokens in November 2022 to $0.07 per million tokens by October 2024, a reduction of more than 280 times in about 18 months. It, drawing on analysis by Epoch AI, found that LLM inference prices had fallen between nine and 900 times per year depending on the task.
The OECD (Organisation for Economic Co-operation and Development), an international policy and economic research organisation, found a similar trend and said that its July 2026 analysis estimated that quality-adjusted prices for text-to-text AI models fell by nearly 80% between January 2024 and April 2026.
The nature of the work being given to AI is changing at the same time. A chatbot might receive a question, process it and return an answer. An AI agent can break a task into several steps, reason through them, call external tools or systems, inspect the results, revise its approach and continue until the task is completed. Each additional step can involve further model calls and more tokens.
The OECD notes that AI agents consume substantially more tokens per task and can increase model-use intensity by orders of magnitude.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.thehindu.com — the content belongs to The Hindu.