Why AI is making economics an architectural discipline.
Most technology leaders have a cloud-bill story. Maybe it was a forgotten development environment or an autoscaling rule that worked a little too well but produced an unexpectedly large cloud bill.However, those experiences taught an entire generation of architects that infrastructure decisions were also economic choices.
While cloud economics focused on the infrastructure beneath the software, AI economics is embedded in the software’s behaviour. Every decision – which model answers a request, how much context it receives, how many tools it calls, how long it reasons, whether another agent reviews its work, when a human must intervene – shapes both the user’sexperience and the economics of the application. Every interaction can set off a chain of model calls, tool invocations and reasoning steps that directly affect both performance and cost.
But the financial impact is only part of the story. The real shift is architectural.

From Cost per Request to Cost per Result
The first instinct is to manage AI costs by counting tokens. Token consumption matters, especially as vendors move toward usage-based models. But cost per token (or even cost per request) is an incomplete measure of economic efficiency.
A low-cost model that produces unreliable answers, triggers retries or requires human escalation may ultimately cost more than a capable model that resolves the task correctly the first time. Likewise, an agent that fails to complete a workflow consumes resources without producing value. The more useful economic unit is therefore the result: cost per resolved customer issue, accepted code change or processed claim. Gartner describes this as moving from cost-per-request to cost-per-result architectures for agentic workflows. The objective is not to minimize computation, but to maximize useful intelligence.
AI economics is the emerging discipline of designing intelligent systems that maximize business outcomes while balancing quality, speed, reliability and cost.
Why Cheaper Intelligence May Create More Expensive Systems
Model prices will almost certainly continue to fall. Gartner forecasts that by 2030, inference on a one-trillion-parameter model could cost providers more than 90% less than it did in 2025. Yet the same research warns that agentic models may consume five to 30 times more tokens per task than a standard generative AI chatbot. Cheaper tokens do not necessarily produce cheaper systems when organizations consume dramatically more of them.
This matters because AI is moving beyond copilots into customer-facing products and autonomous workflows. A few thousand interactions can hide inefficiency; millions amplify every unnecessary reasoning step, bloated context window and redundant tool call. Lower model prices do not eliminate the need for better architecture. They make it more important.
The opportunity, therefore, is not simply to buy cheaper intelligence, but to use intelligence more deliberately.
Context Engineering: Designing Intelligence More Efficiently
If the goal is to maximize useful intelligence while minimizing unnecessary computation, context engineering becomes one of the architect’s most powerful tools. Prompt engineering asks how to phrase instruction, so a model responds well. Context engineering asks a larger systems question: what information, tools, history, rules and state should be available to the model now it makes a decision?
Anthropic describes context as a finite resource and context engineering as the practice of curating the smallest set of high-signal information likely to produce the desired outcome. That principle is important because more context is not automatically better. Sending entire repositories, lengthy conversation histories or every available policy document can increase latency and cost while distracting the model with irrelevant information.
Good context engineering determines what should be retrieved, summarized, cached or omitted, when deterministic software should replace AI, and how long-running agents preserve only the state that still matters.
Prompt engineering optimized conversations. Context engineering optimizes systemsand, by extension, their economics.
AI Economics in Practice
Consider a software company deploying an AI coding assistant across a large engineering organization. In its first version, every request is routed to the most capable model. When a developer asks for an explanation of a method, the assistant sends dozens of repository files, includes a long interaction history and invokes a second model to critique the response. For complex tasks, it launches several agents to explore the codebase in parallel.
At pilot scale, the experience is impressive and the cost manageable. At enterprise scale, the architecture becomes the business model: the same expansive context, premium model and review cycle are repeated for every task.
The team redesigns the system around the outcome. Routine explanations move to smaller models, retrieval follows code dependencies, stable instructions are cached, deterministic tools handle syntax checksand specialist reasoning is reserved for genuinely complex work.
The redesigned system is not simply cheaper. It is faster for routine work, more predictable and easier to govern. It spends more only when additional intelligence is likely to create additional value. That is the difference between managing tokens and practicing AI economics.
Better AI Starts with Better Architecture
The industry is already exposing these tradeoffs as architectural controls. OpenAI offers prompt caching to reduce the cost and latency of reused context and discounted batch processing for work that does not require an immediate response. Google’s experimental Model Optimizer can route requests according to an organization’s preference for cost, quality or a balance of both. Anthropic advises teams to begin with simple, composable patterns and add agentic complexity only when it demonstrably improves the result.Progress is applying similar principles through capabilities such as AI observability and agent tooling that help teams understand, evaluate and optimize AI systems in production rather than relying on intuition alone.
These capabilities point to an important change in mindset. The central design question is no longer “Which is the best model?” but “How much intelligence does this task require, and what is the least complex architecture capable of delivering the result reliably?” The answer shapes model routing, caching, retrieval, evaluation, workflow design and human escalation.
The best AI system will not always use the smallest model or the fewest tokens. It will use resources deliberately.
Measurement Closes the Loop
Economic optimization cannot end at design time. Teams need to see which workflows consume the most resources, where agents repeat work, which models deliver the best result and whether additional spending improves quality.
This is where AI observability becomes essential; not as the solution to AI economics, but as its measurement layer. Traditional monitoring shows whether an application is available. AI systems also require visibility into context, model calls, tool use, latency, quality, retries and cost across an entire workflow. Without that evidence, optimization is guesswork; with it, teams can continuously refine the architecture around real outcomes.
The Next Era of Software Architecture
For nearly three decades, software architecture has focused on helping applications scale. The next era will focus on helping intelligence scale.
Performance, reliability and security will remain essential. But architects must also design systems that balance context, model selection, orchestration, evaluation and governance to deliver the greatest value from every unit of intelligence.
AI has not simply introduced a new technology stack. It has introduced a new design constraint. Every architectural decision now carries an economic consequence. As intelligent systems become the software people rely on every day, the architects who succeed will be those who design not only for performance and reliability, but for economically efficient intelligence.
Authored by Sara Faatz, Senior Director, Awareness at Progress Software
