Garden of Thoughts | June 2026
In AI Engineering, Chip Huyen drops a stat that's hard to shake: one NVIDIA H100 running at peak for a year consumes ~7,000 kWh. Your household uses around 10,000 kWh. So congratulations — your AI assistant is almost as power-hungry as you are.
We talk about electricity like it's the main constraint. It's not — heat is. A chip's TDP (Thermal Design Power) is basically its sweat rating: the maximum heat a cooling system needs to handle under normal load. Pack a few hundred H100s together and suddenly you're less worried about the power bill and more worried about whether the building can exhale fast enough.
Training gets all the attention — the dramatic, expensive, one-time runs that birth a model. But inference is the quiet engine that never turns off. Every autocomplete, every chatbot reply, every API call. It runs on the same hungry hardware, just continuously. Training is the wedding. Inference is the mortgage.
Newer chips are genuinely more efficient per watt. And yet somehow the total power consumption keeps climbing. This is Jevons paradox in action: make something cheaper to run, and people run more of it. We're not escaping the energy problem — we're just climbing it faster with better gear.
Cold air is cheap cooling. Hydro is cheap power. Hyperscalers building in Iceland, Norway, and Quebec aren't chasing scenery — they're chasing physics. Geography is the new moat.
The bottleneck isn't smarter models. It's whether the grid can keep up.
Further reading:
- AI Engineering — Chip Huyen (O'Reilly)
- IEA — Electricity 2025 Report
- Wikipedia — Jevons Paradox
#AIInfrastructure #Energy #Inference #Chips #DataCenters
— Sherry | SherryAnalytics
Curious about data, infrastructure, and where engineering meets the real world.
This post was assisted by Claude (Anthropic).