Inference, not training, is where the agentic era will be won or lost

Jollydeep Kaur explains why inference — not training — is the real cost of AI, how agents multiply token consumption versus chatbots, and why data locality is now a compliance issue.

Jollydeep Kaur, Head – Marketing at NxtGen Cloud Technologies

At the 27th ET Edge CIO&Leader Annual Conference, Jollydeep Kaur, Head – Marketing at NxtGen Cloud Technologies, opened the session by separating two conversations that often get conflated. Training large models, she said, is a game played almost exclusively by a handful of big AI labs with the capital and infrastructure to do it.

Inference is different. It’s the cost enterprises own, and it behaves nothing like a one-time investment. The moment an AI initiative moves from pilot to production, its inference cost becomes a recurring, compounding line item in the company’s P&L. For CIOs used to budgeting technology as a capital expense, this recurring nature is a structural shift in how AI needs to be financially planned for.

Why agents cost more than chatbots

The core of Kaur’s argument rested on a distinction between how chatbots and agents consume compute. A chatbot, she explained, essentially processes one command at a time. Agents, given a task, break it into a business goal, identify the tools and data sources it needs, cross-verify information, retry failed steps, and only then produces an outcome.

Every one of these intermediate steps consumes tokens as if it were its own prompt. Where a chatbot might use one unit of inference to answer a question, an agent completing a comparable task could run through dozens of internal steps to get there. Kaur called the point at which this multiplication starts to outpace linear growth the “crossover point” — the moment inference demand goes exponential, and where many enterprises will first feel the true cost of going agentic.

Compliance adds another layer of cost

Kaur extended the economics argument into governance, noting that when agents operate on regulated data, data locality moves from being a technical or performance consideration to a legal obligation. For enterprises in regulated sectors, this means inference costs cannot be optimised in isolation, decisions about where agents run and how data is processed have to account for compliance from the outset, adding another variable to an already compounding cost equation.

The CIO’s real planning problem

Kaur’s closing point was less about token economics than about mindset: enterprises that plan AI investment purely around capability, without modelling what sustained inference usage will cost as adoption scales, are likely to be caught off guard. Her session left CIOs with a clear takeaway — the success of an agentic deployment isn’t measured only by what the agent can do, but by whether the organisation can sustainably afford to keep it running.

Share on