Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens

· Source: VentureBeat · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cloud Computing & IT Infrastructure · Depth: Advanced, short

Summary

Weka has launched its NeuralMesh 6 software platform and Wekapod 3 hardware line, designed to reduce GPU memory load and inference costs in production AI. The platform introduces Augmented Memory Grid, which aggregates NAND flash to function like GPU memory at a lower cost, addressing the issue of repeated recomputation of pre-calculated tokens in long context windows. NeuralMesh 6 offers composable and virtual multi-tenancy, supporting over 1,000 tenants per cluster and up to 50,000 tenants with 50 composable clusters. It also provides unified file and object storage, claiming two orders of magnitude higher performance than conventional S3 for non-AWS GPU clouds. Other features include metadata-first replication for faster data availability and AlloyFlash, which mixes TLC and QLC NAND flash for optimized cost and performance, alongside Always-On data reduction. This aims to improve GPU utilization and accelerate AI workload deployment.

Key takeaway

For AI Architects and MLOps Engineers scaling large language model inference, you should evaluate Weka's NeuralMesh 6 and Wekapod 3 to mitigate escalating GPU memory costs. By caching 100% of pre-calculated tokens on cheaper flash storage, you can achieve significantly higher GPU utilization and faster deployment of new AI workloads. Consider its composable multi-tenancy and unified storage capabilities to optimize your infrastructure and reduce operational overhead, especially for multi-turn conversational AI or retrieval systems.

Key insights

Extending GPU memory with cheaper flash storage significantly reduces AI inference costs and improves utilization.

Principles

Method

NeuralMesh 6 uses Augmented Memory Grid to cache pre-calculated tokens on NAND flash, routing latency-sensitive tasks to TLC and bulk to QLC, accessible via unified file/object paths.

In practice

Topics

Best for: CTO, VP of Engineering/Data, AI Engineer, AI Architect, MLOps Engineer, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by VentureBeat.