New

Token Economics: Why Energy, Carbon and Cost Must Shape AI Design

Every token your AI generates has a price tag. Three, actually.

There is the bill — whether that arrives from Anthropic, OpenAI, Google, your hyperscaler, or your own inference infrastructure. There is the energy drawn from the grid to produce it. And there is the carbon released into the atmosphere because of it.

For a long time, the working assumption was that engineering teams at least had the first one under control — cost would be optimised, and the other two could follow later. The age of agents has dismantled that assumption.


The problem: token economics has broken

A single user request no longer maps to a single inference call. It triggers chains of model calls, tool invocations, retries, reflection loops, and self-correction passes. A “simple” agentic task can quietly consume thousands of tokens. Multiply that by enterprise scale, and the economics of intelligence stop behaving the way finance teams modelled them last year.

Three forces are compounding at the same time:

Inference has overtaken training as the dominant cost. Training a frontier model is a spike. Serving it to millions of users, every day, for years, is the long tail — and the long tail is now bigger than the spike.

Agents multiply token consumption by an order of magnitude. One user prompt, ten model calls. The same task, ten times the footprint.

Hardware refresh cycles are accelerating. Each new generation of accelerators delivers more performance per watt, but the embodied carbon and water of manufacturing them does not vanish.

The result: cost is increasingly hard to contain — not because engineering teams have stopped optimising, but because the compounding nature of agentic AI workloads, multi-model dependencies, and unpredictable inference patterns is outrunning traditional cost controls. Energy is following the same curve. Carbon is barely measured. Water is rarely mentioned at all.


Why this matters now

We are entering an era where the unit economics of intelligence — cost per token, energy per token, carbon per token, and increasingly water per token — will define which AI products scale and which ones quietly disappear.

A few signals from the past year worth holding in mind:

  • For popular AI services, inference can account for roughly 60 percent of total lifetime energy consumption — and that share grows as user bases scale.
  • Industry analyses project that by 2030, around 70 percent of data centre demand will come from AI inference workloads.
  • Inference data centre capacity is projected to grow from around 2 gigawatts in 2024 to more than 50 gigawatts by 2030.
  • Global data centre electricity consumption is on track to roughly double between 2022 and 2026, with AI workloads as the primary driver.

These numbers point to a simple conclusion: the cost line and the carbon line have converged. The largest lever on AI cost is the same lever as the largest lever on AI sustainability. They are not in tension. They are the same problem, viewed from different angles.


The solution: green software, redefined for the AI era

For years, green software was treated primarily as a carbon conversation. Carbon remains foundational — and the industry has built real standards, real metrics, and real practice around measuring and reducing it. The next chapter is to widen the lens.

Alongside carbon, the resources that quietly shape a workload’s footprint deserve equal engineering attention: the energy it draws, the water consumed in cooling the systems it runs on, and the hardware lifecycle behind every server, chip, and device in the chain.

These dimensions move together. A workload optimised for carbon, running on hardware refreshed every eighteen months, in a data centre drawing heavily from a water-stressed region, has only solved part of the problem. The honest view of green software is multi-dimensional: carbon, energy, water, and waste — considered together across every layer of the stack.

That is the shift worth internalising. From a single-metric mindset to a multi-dimensional one. From carbon accounting to resource intelligence. From sustainability as a reporting exercise to sustainability as an architectural property.


The strategy: five moves that change the math

Token economics, viewed through this lens, becomes a design discipline rather than a billing problem. Five principles are emerging across the leading AI-native organisations:

1. Right-size the model. The largest model is rarely the right model. Matching model capability to task complexity is the single highest-leverage decision in AI cost and footprint. Route simple tasks to smaller models. Reserve frontier capacity for problems that genuinely need it.

2. Cache aggressively, retrieve intelligently. Every token that does not need to be generated is a token that costs nothing, draws nothing, emits nothing. Prompt caching, semantic deduplication, and well-architected retrieval are no longer optimisations — they are the default.

3. Schedule with awareness. Workloads that can run when the grid is cleaner, should. Carbon-aware and energy-aware orchestration has moved from experimental to operational at the leading hyperscalers.

4. Measure across the stack. Token costs at the application layer are shaped by choices at the silicon, infrastructure, platform, and data layers. Optimising one layer in isolation leaves value on the table — and obscures where the real footprint sits.

5. Treat efficiency as an outcome, not a tax. Lower carbon, lower energy, lower water, and lower waste are not concessions to sustainability goals. They are the operational signature of a well-engineered AI system. The teams that internalise this build cheaper, faster, and cleaner products by default.


Why this should be on every executive’s agenda

The companies that will lead the next decade of AI are not the ones spending the most on tokens. They are the ones extracting the most intelligence per resource consumed.

This is not an abstract sustainability narrative. It is the next frontier of competitive advantage, and the leaders building the foundations of this industry already see it clearly:

  • Chip designers are racing to deliver more performance per watt because performance per watt now sets the ceiling on how much intelligence can be deployed at all.
  • Model providers are investing heavily in efficiency — smaller models, smarter architectures, better caching — because the economics of serving billions of queries demand it.
  • Hyperscalers are siting capacity around clean power, water availability, and grid headroom because those constraints now sit on the critical path of growth.
  • Enterprises deploying AI at scale are discovering that token discipline is the difference between AI products that scale profitably and ones that stall at pilot.

If you are leading engineering, architecture, or AI strategy in your organisation, the question worth asking this quarter is not “how do we reduce our AI bill?” The better question is:

Do we understand the full cost of every token we generate — across cost, energy, carbon, water and waste — and are we engineering accordingly?

The organisations that answer this question first will not just be the most sustainable. They will be the most competitive.

The next decade of AI will not be won by whoever spends the most on compute. It will be won by whoever wastes the least.

Token economics is the scoreboard. Green software is the playbook. The leaders who internalise both — across cost, energy, carbon, water, and waste — will define the next era of intelligence.


Technology Bytes is a newsletter on the intersection of emerging technology, sustainability, and enterprise strategy. Subscribe to get the next edition in your feed.