Partnering with Etched: Building the Inference Machine
Summary
Etched, a hardware startup founded in 2022, is developing frontier clusters for AI inference, aiming to maximize intelligence per flop. The company, which will ship its production-ready custom silicon in 2026, has pioneered research breakthroughs like low voltage inference and cluster-scale memory. These innovations enable its inference system to excel in throughput and interactivity across various frontier models, including sparse MoEs, dense transformers, and Mamba. Etched recently taped out its first-generation chip at TSMC, becoming the first post-ChatGPT-era company with a successful full-reticle A0 chip tape-out on leading-edge nodes. Within 40 days, their team brought up the first cluster, achieving Pareto dominant performance on industry-standard benchmarks. Sequoia Capital is leading Etched's \$300 Million Series C funding round at a \$10 Billion pre-money valuation, joined by Jane Street, Andreessen Horowitz, Diffusion, and SK Hynix.
Key takeaway
For Directors of AI/ML evaluating future inference infrastructure, Etched's rapid progress and custom silicon offer a compelling alternative to general-purpose hardware. You should consider how specialized, cluster-scale systems like Etched's, designed for frontier models, could significantly improve your throughput and interactivity metrics. This shift towards purpose-built inference machines suggests a need to re-evaluate your long-term compute strategy, potentially moving beyond off-the-shelf solutions to optimize for cost and performance.
Key insights
Etched is building production-ready, cluster-scale custom silicon for AI inference, achieving superior throughput and interactivity.
Principles
- Design for cluster-scale compute, not just chips.
- Iterate hardware fast to keep pace with model evolution.
- "Production is the product" mantra drives execution.
Method
Etched's approach involves pioneering low voltage inference and cluster-scale memory, then rapidly taping out custom silicon and bringing up clusters to run frontier AI models.
In practice
- Explore low voltage inference for efficiency gains.
- Prioritize cluster-scale memory in system design.
- Co-locate with suppliers to expedite hardware testing.
Topics
- AI Inference Hardware
- Custom Silicon
- Cluster Computing
- Low Voltage Inference
- Frontier AI Models
- Venture Capital Funding
Best for: Investor, AI Hardware Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Sequoia Capital.