The constraint has moved
For several years the binding constraint on AI expansion was accelerator supply. Capital was abundant, demand was insatiable, and the queue for chips defined the pace of the industry.
That constraint has loosened while another has tightened. In several major markets, the limiting factor for new data centre capacity is now the grid connection queue, measured in years rather than quarters.
Operators are responding with on-site generation, long-term power purchase agreements and site selection driven primarily by energy availability. Location decisions that once optimised for latency now optimise for megawatts.
Efficiency is the highest-leverage research
When energy is the constraint, efficiency research stops being a cost-optimisation footnote and becomes the primary route to capability growth. Quantisation, sparsity, distillation, caching and better serving architectures all deliver compounding returns.
The practical evidence is visible in pricing: cost per unit of comparable capability has fallen steeply, driven far more by serving efficiency than by hardware generation improvements.
For engineering teams this is good news. The gains are accessible without frontier research budgets, and most organisations have not yet captured the easy portion of them.
Inference now dominates
Training runs attract attention, but for commercially deployed systems inference has become the dominant consumer of compute by a wide margin, and the gap widens with every successful product.
This changes optimisation priorities. Context caching, response length discipline, model right-sizing per task and aggressive use of small models for routine classification all matter more to a production budget than anything about training.
It also changes the environmental conversation. Reporting that focuses solely on training footprint increasingly misrepresents where the energy actually goes.
Capability convergence and competitive shift
At the frontier, capability differences between leading models are narrowing on most practical tasks. Where a clear gap remains, it tends to be in specific domains rather than across the board.
As capability converges, competition shifts to cost, latency, integration surface and regional availability. That favours buyers and pressures providers to build stickiness through platform features rather than raw model advantage.
The strategic advice for operators is consistent: preserve portability, benchmark on your own workloads, and renegotiate annually.
What could invalidate this outlook
A genuine architectural breakthrough that improves sample efficiency by an order of magnitude would reset every assumption here, and such a result is neither predictable nor impossible.
A faster-than-expected resolution of grid constraints in major markets would similarly change the picture, though the physical timelines involved make dramatic acceleration unlikely within eighteen months.
Planning guidance
Build cost models that assume continued price declines but not capability step-changes. Design systems that can swap models without rewriting product logic.
And invest in the unglamorous efficiency work now: it pays immediately and it hedges against every scenario above.
Comments (0)
Discussion is opening soon. Be the first to comment.