Networking
AI networking encompasses the high-bandwidth, low-latency interconnects — high-bandwidth GPU interconnect, InfiniBand, and 400G Ethernet — that interconnect GPU clusters for distributed training and inference, forming the critical network fabric of AI infrastructure.
AI infrastructure uses a three-tier interconnect hierarchy: high-bandwidth GPU interconnect (within a single GPU node, 900 GB/s), InfiniBand or 400G RoCE (between nodes in a cluster, 200–400 Gb/s), and high-speed Ethernet or optical (between data centers, 100–400 Gb/s). Each tier serves a different communication pattern in the AI training process.
high-bandwidth GPU interconnect enables 8 GPUs within a node to share memory directly at 900 GB/s — critical for model parallelism where a single model is split across GPUs. InfiniBand connects nodes in a cluster for data parallelism, where gradient updates are synchronised across nodes after each training step. Both tiers must be non-blocking to maintain training efficiency.
InfiniBand is the standard for AI clusters, offering 200 Gb/s (HDR) or 400 Gb/s (NDR) bandwidth with 1–2 microsecond latency. It includes RDMA (Remote Direct Memory Access), allowing GPUs in different nodes to communicate directly without CPU involvement — critical for distributed training performance.
Ethernet (400G RoCE — RDMA over Converged Ethernet) is a lower-cost alternative that provides similar RDMA capabilities over standard Ethernet. However, Ethernet has higher latency (5–10 microseconds) and more congestion under heavy load. Frontier-scale AI clusters universally use InfiniBand; cost-optimised clusters may use RoCE.
AI clusters use fat-tree network topology — a hierarchical arrangement of switches where bandwidth increases toward the root, ensuring non-blocking communication between any pair of GPUs. A 2-tier fat-tree supports up to ~1,000 GPUs; a 3-tier fat-tree supports 10,000+. The cost of switching scales with cluster size, representing 8–12% of total cluster capital.
Constellation's GPU clusters use full-fat-tree InfiniBand topologies, ensuring any GPU can communicate with any other GPU at full bandwidth — a non-negotiable for efficient distributed training.
Learn how you can participate in the AI infrastructure investment opportunity across the UAE, GCC, and India.
Invest with CAT