NewsProject Update02 May 2026

Inference Is the New Growth Engine

Training built the AI market. Inference is what will scale it, and it is rewriting the requirements for the data centres underneath.

AI Infrastructure · Compute Markets

Quick actions

Read time

4 min · 792 words

Category

Project Update

Embargo

Lifted · Public

There is a quiet but consequential shift underway in how artificial intelligence consumes compute. For the first wave of the AI boom, the headline number was training: the eye-watering cost of building ever-larger frontier models in vast, purpose-built clusters. That work continues, but it is no longer where most of the growth lives. The market is moving from a training economy (building the brain) to an inference economy (using it, continuously, at scale).

The figures behind that shift are stark. Inference workloads are projected to account for roughly two-thirds of all AI compute in 2026, up from about one-third in 2023 and half in 2025.¹ The global AI inference market is forecast to grow from around USD 106 billion in 2025 to roughly USD 255 billion by 2030, a compound annual growth rate near 19 percent.² By the end of the decade, some analysts expect the inference market to dwarf training by an order of magnitude.

~2/3

of all AI compute is expected to be inference in 2026, up from one-third in 2023.

Why inference changes the economics

The reason is structural. A model is trained once, but it is queried indefinitely. Every chatbot response, every code completion, every fraud check and every recommendation is an inference call, and as AI is embedded into mainstream products, those calls multiply without end. Industry analysis suggests inference can represent 80 to 90 percent of the lifetime cost of a production AI system, precisely because it runs around the clock.¹ Training is a capital event; inference is an operating reality.

That distinction matters enormously for infrastructure. Training is bursty, tolerant of latency, and can be concentrated in a handful of mega-clusters wherever power is cheapest. Inference is the opposite: it is continuous, latency-sensitive and demand-following. It needs to sit close to users, run reliably every hour of every day, and scale in step with adoption. In power terms, the divergence is already visible: inference demand is forecast to grow at around 35 percent annually to exceed 90 gigawatts by 2030, outpacing training, which is expected to reach more than 60 gigawatts over the same period.³

What inference-first infrastructure looks like

Meeting continuous, high-volume inference demand places specific requirements on an AI data centre. Rack densities are climbing sharply as GPU platforms advance, pushing well beyond what air cooling can economically handle. Direct-to-chip and immersion liquid cooling are becoming the default for high-density AI halls, not an exotic option. Electrical topology has to be both resilient and modular, able to add capacity in step with demand rather than in a single, speculative bet. And because inference workloads run constantly, operational reliability, the difference between three nines and five nines of availability, translates directly into customer economics.

There is also a premium on flexibility. The hardware underneath inference is evolving rapidly: each new GPU generation shifts the power-per-rack and thermal envelope that facilities must accommodate. A data centre commissioned today may host two or three hardware generations over its early life. Designing for adaptability (flexible cooling deployment, scalable power distribution, white space that can be reconfigured) is what separates infrastructure that ages gracefully from infrastructure that is obsolete on arrival.

Positioning for the inference economy

This is the environment INSITE DC, an Australian data centre developer, has built its platform for. Our Melbourne data centre campus is engineered for high-density, liquid- and immersion-cooled AI workloads spanning both training and inference, with a scalable electrical architecture and a target PUE of circa 1.3. It is a GPU data centre aligned to next-generation accelerator platforms and designed with modularity at its core so that capacity and cooling can evolve as the workloads do. Our co-design philosophy (aligning with partners on density, cooling and network topology before the scope hardens) is particularly suited to inference customers, whose deployments grow and change continuously rather than landing all at once.

The training era proved what AI could do. The inference era is where that capability meets the real world, billions of queries at a time. For data centre operators, it is the more demanding test (continuous, latency-bound and relentlessly cost-sensitive), and it is the one that will define the next decade of AI infrastructure. The growth engine has changed gears. The infrastructure has to change with it.

Sources

  1. 1.Introl: AI Inference vs Training Infrastructure Economics
  2. 2.MarketsandMarkets: AI Inference Market Size, Share & Growth, 2025 to 2030
  3. 3.McKinsey: The next big shifts in AI workloads and hyperscaler strategies
  4. 4.Grand View Research: AI Inference Market Size and Trends, 2030

This article is provided for general information and thought-leadership purposes. Market figures are drawn from third-party research as cited and are indicative; capacity, timing and design figures relating to INSITE DC reflect current development plans and are subject to change.

About INSITE DC

INSITE DC is an Australian developer of next-generation AI and hyperscale data centre infrastructure. Our flagship Melbourne campus is engineered from the ground up for high-density GPU compute, liquid and immersion cooling, and a target PUE of circa 1.3, with a development pathway scaling beyond 400MW and a 1GW+ pipeline across Australasia. Built on our values of Insight, Never-Fail Reliability, Service Excellence, Integrity, Trusted Partnership and Environmental Responsibility, we co-design infrastructure with our partners rather than for them.