NVIDIA NVLink: The Scale-Up Network for AI Factories

· Source: NVIDIA Technical Blog · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cloud Computing & IT Infrastructure · Depth: Advanced, long

Summary

NVIDIA NVLink is presented as the purpose-built scale-up networking fabric for AI factories, designed to accelerate AI inference, training, and parallel computing workloads requiring extensive GPU-to-GPU communication. The Sixth Generation NVLink interconnect with NVLink 6 Switch offers 3.6 TB/s bidirectional GPU-to-GPU bandwidth and 260 TB/s rack-level bandwidth in a 72-GPU domain, achieving 3X lower end-to-end latency and 10X higher packet rates than off-the-shelf Ethernet. It also supports SHARP in-network compute (130 TFLOPS) for collective operations and includes rack-level resiliency features. This co-designed technology, part of NVIDIA's annual AI infrastructure roadmap, has demonstrated up to 2.3X higher decode throughput for MoE models like DeepSeek-R1 and Qwen 235B, and a 50X improvement in MoE inference performance per watt from Hopper to Blackwell platforms. NVLink-C2C extends this with 1.8 TB/s coherent CPU-GPU bandwidth, 7x PCIe Gen6, while NVLink Fusion allows custom silicon integration.

Key takeaway

For AI Architects and MLOps Engineers designing or upgrading AI factories, prioritizing a purpose-built scale-up networking fabric like NVIDIA NVLink is crucial. You should evaluate solutions not just on raw bandwidth, but on delivered full-system performance, end-to-end latency, in-network compute, and integrated resiliency features. Opting for a mature, co-designed stack can significantly reduce deployment risk and maximize your factory's sustained ROI by improving utilization and lowering cost-per-token.

Key insights

Purpose-built scale-up networking like NVLink is critical for AI factory performance, resiliency, and economic efficiency.

Principles

In practice

Topics

Best for: CTO, VP of Engineering/Data, Director of AI/ML, AI Architect, MLOps Engineer, AI Hardware Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by NVIDIA Technical Blog.