COMPUTEX 2026 | NVIDIA Keynote | Extreme Co-Design: Building the AI Factory

· Source: NVIDIA · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cloud Computing & IT Infrastructure, Robotics & Autonomous Systems · Depth: Advanced, long

Summary

NVIDIA's SVP Networking, Kevin Deierling, presented "Extreme Co-Design: Building the AI Factory" at Computex 2026, highlighting three major transformations: the shift to accelerated computing, the rise of generative AI and agentic reasoning, and the re-architecture of data centers into AI factories. With Moore's Law ending and AI model parameters growing 10x annually, NVIDIA focuses on driving down cost per token. This involves optimizing models from FP16 to NVFP4 and redesigning data centers as "AI factories" across a five-layer stack. Key innovations include smoothing energy demand (DSX Flex, DSX LPS) to deploy 40% more GPUs per gigawatt, co-designing seven chips like the Rubin GPU and Vera CPU for 10x token improvement, and implementing liquid cooling and co-packaged optics for infrastructure. The emergence of AI agents necessitates new AI storage architectures (CMX, STX) for 5x faster token throughput, alongside optimized models and CUDA-X accelerated applications.

Key takeaway

For AI Architects and MLOps Engineers scaling AI infrastructure, NVIDIA's "Extreme Co-Design" framework offers a blueprint for maximizing efficiency. You should evaluate your data center's energy management, chip integration, and cooling solutions, considering liquid cooling and co-packaged optics. Prioritize specialized AI storage (CMX, STX) and agentic frameworks to optimize token throughput and prepare for future AI workloads, ensuring your AI factory generates maximum revenue.

Key insights

NVIDIA's "Extreme Co-Design" optimizes AI factories across a five-layer stack, from energy to applications, to meet escalating AI demand efficiently.

Principles

Method

NVIDIA's "five-layer cake" co-design approach optimizes energy, chips, infrastructure, models, and applications, from bottom-up and top-down, to maximize tokens per watt/dollar.

In practice

Topics

Best for: Investor, CTO, AI Engineer, AI Architect, Director of AI/ML, MLOps Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by NVIDIA.