AI’s Biggest Bottleneck Isn’t GPUs Anymore

· Source: Towards AI - Medium · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cloud Computing & IT Infrastructure · Depth: Intermediate, long

Summary

SK Hynix recently completed the largest foreign IPO in US history, raising \$26.5 billion on Nasdaq, while NVIDIA's stock declined 15% from its May peak and Micron's tripled. This shift highlights a critical change in the AI infrastructure bottleneck, moving from GPUs to High Bandwidth Memory (HBM). H100 GPU rental prices have fallen from \$3.20/hour to under \$2.60/hour, contrasting sharply with a tenfold rise in DRAM spot prices since summer 2025. Major tech companies like Meta, OpenAI, Anthropic, Google, Amazon, and Microsoft are developing custom AI silicon, further diversifying the compute landscape. Concurrently, China has established its first 100,000-card domestic AI cluster, the Sugon 8000, using Hygon accelerators. This indicates AI chips are commoditizing, memory is the new center of power, and the global AI supply chain is splitting into parallel US-aligned and fully domestic Chinese tracks.

Key takeaway

For Directors of AI/ML or CTOs planning future infrastructure, recognize that your primary hardware cost driver is shifting from GPUs to High Bandwidth Memory. You should prioritize optimizing memory utilization and explore custom silicon options to reduce dependency on single GPU vendors. This diversification will lower inference costs and allow you to select hardware tailored to specific workloads, moving towards a more resilient and cost-effective AI compute strategy.

Key insights

AI's core bottleneck has moved from GPUs to High Bandwidth Memory, reshaping the industry's power dynamics.

Principles

In practice

Topics

Best for: VP of Engineering/Data, Executive, Entrepreneur, Director of AI/ML, CTO, Investor

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Towards AI - Medium.