NVIDIA NVLink: The Scale-Up Network for AI Factories
Summary
NVIDIA NVLink is presented as the purpose-built scale-up networking fabric for AI factories, designed to accelerate AI inference, training, and parallel computing workloads requiring extensive GPU-to-GPU communication. The Sixth Generation NVLink interconnect with NVLink 6 Switch offers 3.6 TB/s bidirectional GPU-to-GPU bandwidth and 260 TB/s rack-level bandwidth in a 72-GPU domain, achieving 3X lower end-to-end latency and 10X higher packet rates than off-the-shelf Ethernet. It also supports SHARP in-network compute (130 TFLOPS) for collective operations and includes rack-level resiliency features. This co-designed technology, part of NVIDIA's annual AI infrastructure roadmap, has demonstrated up to 2.3X higher decode throughput for MoE models like DeepSeek-R1 and Qwen 235B, and a 50X improvement in MoE inference performance per watt from Hopper to Blackwell platforms. NVLink-C2C extends this with 1.8 TB/s coherent CPU-GPU bandwidth, 7x PCIe Gen6, while NVLink Fusion allows custom silicon integration.
Key takeaway
For AI Architects and MLOps Engineers designing or upgrading AI factories, prioritizing a purpose-built scale-up networking fabric like NVIDIA NVLink is crucial. You should evaluate solutions not just on raw bandwidth, but on delivered full-system performance, end-to-end latency, in-network compute, and integrated resiliency features. Opting for a mature, co-designed stack can significantly reduce deployment risk and maximize your factory's sustained ROI by improving utilization and lowering cost-per-token.
Key insights
Purpose-built scale-up networking like NVLink is critical for AI factory performance, resiliency, and economic efficiency.
Principles
- AI factory performance hinges on all-to-all bandwidth and low latency.
- Extreme co-design across the stack optimizes delivered performance.
- Resiliency features are essential for sustained AI factory goodput.
In practice
- Evaluate scale-up fabrics based on full-system delivered performance.
- Prioritize solutions with integrated software and robust supply chains.
- Consider coherent CPU-GPU connectivity for agentic workloads.
Topics
- NVIDIA NVLink
- AI Factories
- Scale-Up Networking
- GPU-to-GPU Communication
- MoE Models
- Data Center Infrastructure
Best for: CTO, VP of Engineering/Data, Director of AI/ML, AI Architect, MLOps Engineer, AI Hardware Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by NVIDIA Technical Blog.