Day 22 — The Silicon in the Rack
Day 21 opened the embodied carbon arc by drawing the boundary between operational emissions (what the system draws while running) and embodied emissions (what the system cost to exist). That framing named three embodied categories: hardware, training, and facility. Today’s issue takes the first one apart.
The AI accelerator, whether GPU, TPU, or purpose-built inference chip, has become the defining artifact of the AI era. It sits in the rack, draws power, generates heat, and delivers the inference every dashboard measures. It also carries a manufacturing footprint that arrived on the day it was installed, before it served its first token, before it drew its first watt. That footprint was paid in silicon fabrication, in packaging, in the assembly of the server around it, and in the logistics that moved it from a factory to a data center floor. The number is not small, and it is not evenly distributed across the hardware’s life. The choices a team makes about which silicon to buy, how long to run it, and when to replace it determine how that fixed footprint is spread across the work the hardware actually does.
This issue is about that decision: the hardware refresh cycle as an embodied-carbon lever, and the arithmetic that connects it to per-inference emissions.
A cloud-native AI platform team is planning its next hardware refresh. The current fleet, roughly three years old, has served the company’s inference workloads faithfully. A newer generation of accelerators is available with meaningfully better performance per watt. The operational case for the upgrade is clean: fewer chips, lower electricity draw, lower cooling load. The finance case supports it. The refresh is approved.
The sustainability lead asks one question before signing off. What happens to the embodied footprint of the fleet being retired, and what is the embodied footprint of the fleet coming in?
The audit that follows surfaces three things nobody had costed.
The retired fleet’s remaining embodied value. The outgoing accelerators had a rated operational lifespan of five to seven years. They were being retired at three. The embodied footprint of the outgoing hardware had been amortized against roughly half the useful life it could have served. The rest of that footprint, emitted long before retirement and paid for the moment the hardware entered service, would now be carried by whatever secondary market the hardware entered, if any.
The incoming fleet’s embodied cost. The new accelerators had a substantially higher embodied footprint per unit than the outgoing generation. Advanced packaging, larger HBM stacks, and increasingly complex manufacturing all contribute to higher embodied emissions per accelerator. The operational savings were real. The embodied cost was real too. Whether the trade came out positive depended entirely on how long the new fleet would run, and on the carbon intensity of the grid it would run on.
The refresh cadence itself. The team had been on a three-year refresh cycle by default. Nobody had chosen three years explicitly. It had emerged from a mix of depreciation schedules, vendor roadmaps, and internal upgrade momentum. The embodied consequences of that cadence, repeated across every refresh, had never been on the table.
Operational efficiency usually improves with every hardware generation. Embodied carbon often rises with it. Whether the upgrade is actually greener depends less on peak efficiency than on how long the new hardware remains in service, and on the grid it runs on.
The obvious objection is that older hardware is less efficient, and running it longer means burning more electricity per unit of work. That is true, and the trade-off is real. But the framing of “newer is greener” collapses without the embodied side. A generation whose operational savings would take four years to pay back its embodied cost is not greener if it is refreshed in three. The refresh cadence is one variable that decides which side of the trade wins. The grid the fleet runs on is the other.
The Grid Beneath the Chip
The payback arithmetic on any hardware refresh depends on where the fleet is plugged in. Embodied carbon is fixed the day the hardware ships. Operational carbon scales with the grid.
On a coal-heavy grid with high emissions per kilowatt-hour, operational carbon dominates the total footprint. A more efficient chip pays back its embodied debt quickly, often inside a year, and the case for refresh strengthens. On a near-zero-carbon grid running hydro, nuclear, or high renewables, operational carbon is close to negligible. Embodied carbon then represents almost the entire footprint of the fleet, and hardware life extension becomes the only meaningful lever available.
The practical consequence is that the same refresh decision produces opposite recommendations in different regions. A hyperscaler’s carbon-heavy region and its clean-grid region should not be on the same refresh cadence. The greener choice depends on where the sockets are.
Four rungs move the hardware refresh cycle from a procurement default into an embodied-carbon decision.
Measure. The starting point is a number the team may never have asked for: the embodied footprint of the hardware in the rack. Vendors are increasingly publishing this. Product carbon footprint documents, life-cycle assessments, and disclosure statements are now available for major accelerator families. Third-party datasets from Boavizta and similar sources fill the gaps for hardware where vendor data is thin. Pair the embodied figure with the operational figure for the grid the hardware runs on. The number will never be perfect, but it needs to exist as a first-class figure alongside the operational one.
Extend. For most hardware, the largest embodied-carbon lever is time. Extending hardware from three years to five, or five to seven, spreads the same embodied footprint across far more useful work before any operational savings are considered. Extension is not a passive choice. Silicon run near thermal limits for years accumulates real wear: thermal paste degrades, cooling efficiency drifts, failure rates rise, and vendor support windows close. Extending useful life requires proactive maintenance, thermal management, and graceful failover design. Done well, extended-life hardware moves into secondary deployment for lower-demand workloads or into non-production environments where a modest failure rate is acceptable.
Match. Not every workload needs the newest silicon. Modern frontier inference relies on hardware-accelerated features that older generations lack: FP8 and INT4 quantization, transformer engines, native FlashAttention support. Running a current frontier model on a five-year-old architecture may require far more nodes or far longer batch windows, which destroys the operational efficiency the newer chip was supposed to deliver and may make latency SLAs impossible. Older silicon is not obsolete, it is specialized. It runs smaller specialized models (SLMs), embedding generation, retrieval preprocessing, batch scoring, evaluation pipelines, and CI/CD workloads well. Matching workload to hardware generation, running the newest silicon where it earns its embodied cost and older silicon where it still serves, spreads the fleet’s embodied footprint against the workloads best suited to carry it.
Choose. When a refresh is genuinely justified, the choice is not only which chip but which supplier, which manufacturing process, which packaging technology, and which region of manufacture. Embodied footprints vary meaningfully across all four. The vendor with the best operational efficiency is not always the vendor with the lowest embodied cost, and the two numbers together give a more honest picture than either alone. Procurement decisions become disclosure-informed.
The hardware refresh cycle touches every seat at the table, but the embodied dimension has usually been invisible to most of them.
Engineering. Hardware choice shows up in benchmarks, latency budgets, and per-inference cost. It now shows up in embodied carbon per inference too, a number the team can influence with the same rigour it already applies to throughput.
Architecture and CTO. The architecture diagram gains a lifespan dimension. Systems designed around a three-year hardware lifetime look different from systems designed around a seven-year one, in redundancy, in software portability, in graceful degradation, in the willingness to depend on features specific to a single silicon generation.
Platform and Infrastructure. Procurement gains an embodied-carbon column and a grid-carbon column. Vendor conversations extend from performance-per-watt to embodied-per-unit. Refresh cadence becomes an explicit choice with a documented rationale, differentiated by region rather than inherited from a global depreciation schedule.
Sustainability and ESG. Hardware embodied carbon is one of the largest components of the widened AI footprint, and one of the easiest to defend once measured. The reporting boundary now includes a category with real numbers behind it and two clear levers, refresh cadence and grid choice, that leadership can influence.
Business and Product. Total Cost of Ownership was already a hardware-lifecycle conversation. Total Footprint of Ownership (TFO) makes it a longer one, and gives leadership a ready-made vocabulary for board decks and disclosures alongside TCO. A fleet run to seven years on a clean grid costs less per inference on both axes than the same fleet run to three on a carbon-heavy grid, provided the operational efficiency of the older generation still serves the workload.
Five seats, one silicon.
The work begins with an inventory and ends with a cadence.
List every accelerator, every server, and every storage system in the current AI fleet. For each, record three numbers: the embodied footprint per unit (from vendor disclosure or third-party estimate), the date it entered service, and the carbon intensity of the grid it runs on. Estimate the fleet’s annual embodied carbon by allocating each asset’s embodied footprint across its expected service life, and pair it with the annual operational carbon for the same asset. That paired number is the starting point for every subsequent decision.
Pick the largest category, usually accelerators, and ask two questions. Is the current refresh cadence a chosen number or an inherited one? And is that cadence uniform across regions with very different grid intensities? If either answer is “inherited” or “uniform,” propose an alternative. A one-year extension of useful life on a clean-grid region is often the largest single embodied-carbon action available without changing any other decision.
Document the new cadence, the regional differentiation, and the rationale in the procurement policy.
One inventory, one cadence, one rationale. That is the work.
The greenest chip is often the one you already own, running on the cleanest grid you have access to.
Day 23 turns to the second embodied category: the model itself. The training run that produced it, the fine-tunes that shaped it, and the arithmetic that decides how thinly that one-time cost is spread across every inference the model serves.
- ISO/IEC 21031:2024, Software Carbon Intensity (SCI) specification, with M-term for embodied emissions.
- Green Software Foundation. Software Carbon Intensity, measurement guidance. greensoftware.foundation
- Green Software Foundation. Carbon-aware computing and grid intensity APIs. greensoftware.foundation
- Boavizta. BoaviztAPI, lifecycle assessment methodology and tooling for IT equipment. boavizta.org and github.com/Boavizta/boaviztapi
- Gupta, U. et al. (2021). Chasing Carbon: The Elusive Environmental Footprint of Computing. arXiv:2011.02839
- Gupta, U. et al. (2022). ACT: Designing Sustainable Computer Systems with an Architectural Carbon Modeling Tool. arXiv:2205.14226
- NVIDIA. Product Carbon Footprint disclosures for datacenter GPUs, and Transformer Engine / FP8 documentation. nvidia.com/sustainability
- Google Cloud. TPU sustainability and lifecycle disclosures. cloud.google.com/sustainability
- AMD, Intel. Product-level carbon disclosures for datacenter accelerators.
- Electricity Maps. Real-time grid carbon intensity data. electricitymaps.com
- WattTime. Marginal emissions data and grid intensity signals for carbon-aware compute. watttime.org
- Greenhouse Gas Protocol. Corporate Value Chain (Scope 3) Accounting and Reporting Standard. ghgprotocol.org