Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens
Summary
Weka has launched its NeuralMesh 6 software platform and Wekapod 3 hardware line, designed to reduce GPU memory load and inference costs in production AI. The platform introduces Augmented Memory Grid, which aggregates NAND flash to function like GPU memory at a lower cost, addressing the issue of repeated recomputation of pre-calculated tokens in long context windows. NeuralMesh 6 offers composable and virtual multi-tenancy, supporting over 1,000 tenants per cluster and up to 50,000 tenants with 50 composable clusters. It also provides unified file and object storage, claiming two orders of magnitude higher performance than conventional S3 for non-AWS GPU clouds. Other features include metadata-first replication for faster data availability and AlloyFlash, which mixes TLC and QLC NAND flash for optimized cost and performance, alongside Always-On data reduction. This aims to improve GPU utilization and accelerate AI workload deployment.
Key takeaway
For AI Architects and MLOps Engineers scaling large language model inference, you should evaluate Weka's NeuralMesh 6 and Wekapod 3 to mitigate escalating GPU memory costs. By caching 100% of pre-calculated tokens on cheaper flash storage, you can achieve significantly higher GPU utilization and faster deployment of new AI workloads. Consider its composable multi-tenancy and unified storage capabilities to optimize your infrastructure and reduce operational overhead, especially for multi-turn conversational AI or retrieval systems.
Key insights
Extending GPU memory with cheaper flash storage significantly reduces AI inference costs and improves utilization.
Principles
- Cache 100% of pre-calculated tokens to avoid recomputation.
- Mix TLC and QLC flash for optimal cost/performance.
- Unified file/object storage eliminates data duplication.
Method
NeuralMesh 6 uses Augmented Memory Grid to cache pre-calculated tokens on NAND flash, routing latency-sensitive tasks to TLC and bulk to QLC, accessible via unified file/object paths.
In practice
- Implement Augmented Memory Grid for long context AI models.
- Utilize composable multi-tenancy for isolated AI workloads.
- Adopt unified storage for training and inference data.
Topics
- AI Inference Optimization
- GPU Memory Extension
- Flash Storage
- NeuralMesh 6
- Augmented Memory Grid
- Multi-tenancy
Best for: CTO, VP of Engineering/Data, AI Engineer, AI Architect, MLOps Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by VentureBeat.