The production platform for open-weight AI inference
Summary
Together has released a major update to its inference platform, offering users complete control over performance, cost, and quality for open-weight AI models without requiring a custom infrastructure stack. This enhanced platform enables models to go live in minutes, supporting production-grade deployments with features like multiple deployments behind a single stable endpoint, safe changes via canary, blue-green, and rolling updates with auto-rollback, and real-traffic testing through A/B and shadow testing. The service also provides flexible autoscaling across regions and offers optimized deployment profiles, leading to approximately 4x faster warm starts. Additionally, Together is launching a closed beta for custom training, including full-weight and LoRA reinforcement learning and supervised fine-tuning, allowing direct deployment of checkpoints to production. The platform aims to simplify the entire model lifecycle, from training to continuous improvement.
Key takeaway
For MLOps Engineers managing open-weight AI model inference, Together's updated platform simplifies complex deployments and continuous iteration. You can achieve production-grade stability and performance in minutes, leveraging features like safe rollouts with auto-rollback and real-traffic testing. This allows you to maintain complete control over your models' performance, quality, and cost, significantly reducing the operational overhead of building and maintaining custom inference stacks. Consider exploring the custom training beta to streamline your model development lifecycle.
Key insights
The platform provides comprehensive control and simplified deployment for open-weight AI inference, integrating training to production.
Principles
- Open-weight models offer control over performance, quality, and functionality.
- Inference platforms should integrate training, deployment, and continuous improvement.
- Production readiness must be foundational, not an afterthought.
Method
The platform supports deploying models from various sources, selecting optimized profiles, scaling based on diverse metrics, and safely rolling out changes with A/B or shadow testing.
In practice
- Deploy open-weight models from Hugging Face or S3.
- Use optimized deployment profiles for faster setup.
- Implement canary or A/B testing with real traffic.
Topics
- AI Inference Platform
- Open-weight Models
- Model Deployment
- MLOps
- Canary Deployments
- Custom Model Training
Best for: CTO, VP of Engineering/Data, Director of AI/ML, MLOps Engineer, Machine Learning Engineer, AI Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Together AI | The AI Native Cloud - Together.ai.