Growing pains: how distributed AI training changes the network between datacenters
Large-scale AI training has already escaped the confines of a single datacenter.
Google said Gemini was trained synchronously across clusters in multiple locations; Microsoft has connected AI data centres in Wisconsin and Georgia into what it describes as one distributed AI supercomputer; AWS has connected AI compute clusters across wide areas to allow Anthropic to build Claude models; Meta has built high-capacity datacenter interconnects to support model training; and CoreWeave and Google Cloud recently announced cross-cloud training with a private interconnect, with Azure likely to follow later in the year.
This growing geographic spread of model training is partly due to the limits of single datacenters or campus clusters, which can be strained as they try to meet the compute and power demands of new models.
Cisco estimates that training models today can require clusters with tens of thousands of GPUs.
By 2030, the largest individual frontier training runs could draw 4-16GW of power, according to researchers at Epoch AI.
Distributing compute allows hyperscalers and datacenter operators to build in locations with more available power or space and fewer planning constraints.
How to train an LLM At the core of an LLM is a neural network with billions of numerical parameters, or weights, that are adjusted during training.
A batch of training data passes through the model, which makes a prediction.
The system measures the error and calculates how the weights should change.
The weights are updated and the process starts again.
At the scale of modern models, this work is distributed across thousands of GPUs and other accelerators, such as AWS Trainium and Google TPUs.
Different GPUs can process separate batches of data or different parts of the model.
But they cannot work entirely independently: at various points they must exchange results and synchronize before the next stage of training can begin.
What networks carry between the accelerators is generally not the text or images used to train the model, but large arrays of numerical data such as gradients and intermediate results.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.