As AI factories expand from thousands of GPUs to tens of thousands, a slow or congested connection can cause a large number of accelerators to wait, making it difficult to fully utilize the peak performance of GPUs. Today, NVIDIA announced the new Spectrum-6 Ethernet switch system, with a transmission speed of 102.4 Tb per second, double the capacity of the previous generation. NVIDIA claims that the Spectrum-X platform can maintain network efficiency of up to 95% in environments with over 100,000 GPUs deployed. CoreWeave, Microsoft, Nebius, SpaceXAI, and Tesla will be the first adopters. Spectrum-6 is designed for the Vera Rubin platform and AI factories with a scale of 1 GW. As advanced models, agentic AI, and physical AI require more GPUs to work together, the network has evolved from a data transmission tool to a core infrastructure that determines GPU utilization, model training time, and cost per token. AI factories are moving towards a scale of one million kilowatts. GPU peak performance does not equal overall computing power. Traditional enterprise data centers' Ethernet primarily handles north-south traffic between users, servers, and storage systems. However, large-scale AI model training and inference require thousands to tens of thousands of GPUs to continuously exchange data, creating massive east-west traffic. In collective communication tasks, each GPU must synchronously exchange computational results. If some nodes experience delays, packet loss, or network congestion, it may force other GPUs to wait, causing expensive accelerators to fail to maintain full load. Therefore, looking only at the theoretical peak performance of GPUs is insufficient to measure the actual output of AI factories. Switch bandwidth, network routing, congestion control, fault recovery, and optical transmission all affect whether the entire system can scale effectively. Spectrum-6 reaches 102.4 Tb/s, with capacity increased by one fold compared to the previous generation. Spectrum-6 is the core of the new Spectrum-X Ethernet platform, with a single switch system transmission speed of 102.4 Tb per second, a capacity increase of one fold compared to the previous generation, and co-designed with ConnectX-9 SuperNIC, network software, and the Vera Rubin platform. Unlike general-purpose Ethernet, Spectrum-X introduces adaptive routing, advanced congestion control, and real-time telemetry, which can balance traffic among multiple available paths based on network conditions. When some connections fail, the system can also rearrange the transmission path and recover data that cannot reach the destination. NVIDIA states that the AI network performance of Spectrum-X can reach 1.6 times that of general-purpose Ethernet, and even if the deployment environment exceeds 100,000 GPUs, the network efficiency can still be maintained at up to 95%. Through hardware-accelerated multi-plane network topology, the number of switches required for data centers can be reduced by 1.7 times. However, the above numbers are all results provided by NVIDIA based on specific test or design conditions, and actual performance still depends on model architecture, cluster scale, network topology, software settings, and workload. CoreWeave, Microsoft, Nebius, SpaceXAI, and Tesla will be the first companies to adopt Spectrum-6. Among them, CoreWeave, Microsoft, and Nebius will deploy Vera Rubin infrastructure equipped with Spectrum-6, providing developers, startups, and enterprises with access. Min Jun, Network Product Director of CoreWeave, stated that the network is the key to maintaining performance of AI workloads in large-scale environments. The introduction of Spectrum-6 and liquid-cooled Spectrum-X infrastructure will help provide the bandwidth, resilience, and efficiency required for advanced model training and inference. CoreWeave has also adopted the SN6600-LD equipped with Spectrum-6 chips to build the Vera Rubin NVL72 exchange network. This liquid-cooled high-density exchange rack can provide a capacity of 1.64 Pb/s per rack, a 100% increase over the previous generation air-cooled exchange, and connects Vera Rubin GPUs through a non-blocking, multi-plane, multi-track network architecture. Laurelle Roseman, Vice President of Global Partnerships at Nebius, pointed out that at a scale of one million kilowatts, performance depends on whether each GPU can maintain synchronization, 'avoiding a single slow connection slowing down the entire job.' She stated that Spectrum-6 is designed to address this issue, and even as customers' workloads continue to expand, they must maintain high speed and resilience. Spectrum-6 simultaneously supports pluggable optical components and co-packaged optical components, and the product can also adopt liquid cooling, forming an end-to-end cooling architecture with Vera Rubin computing racks. Traditional pluggable optical transceivers need to transmit electrical signals from the switch chip to the optical module on the panel. As transmission speed and switch bandwidth increase, the power consumption and heat dissipation problems caused by signal transmission distance become more pronounced. Co-packaged optics moves optical components closer to the switch chip, shortening the transmission distance of high-speed electrical signals. According to NVIDIA's data, Spectrum-X photonics technology can improve energy efficiency by 5 times compared to pluggable architectures, and the mean time between failures is improved by 10 times. CoreWeave, Lambda, and Oracle's OCI are among the first adopters. For AI factories with tens of thousands of GPUs, the power consumption, heat dissipation, and failure rates of the switches and optical modules themselves are amplified. This also makes CPO, silicon photonics, and liquid cooling technologies upgrade from individual components to key factors affecting the total cost of ownership of data centers. Spectrum-6 also reflects NVIDIA's expansion of its competitive scope from GPUs to network cards, switches, DPUs, optical components, network software, and cooling systems. NVIDIA calls this strategy 'vertical integration, horizontal openness': on the one hand, co-designing chips, systems, and software, and on the other hand, still supporting standard Ethernet, open network operating systems, and different RDMA transmission modes, reducing the integration difficulty for customers when adopting the complete platform. When AI factories reach the scale of one million kilowatts, computing performance is no longer determined by a single chip, but by GPUs, CPUs, networks, optical transmission, power, and cooling systems working together. The strategic significance of Spectrum-6 is not just selling more switches, but allowing NVIDIA to further grasp the technology stack from GPUs to the entire AI factory.
FACT BOX
- Source: PR Times
- Category: New Product
- Organizations: CoreWeave / Microsoft / Nebius
- Products / services: Spectrum-6