The production platform for open-weight AI inference

· Source: Together AI | The AI Native Cloud - Together.ai · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cloud Computing & IT Infrastructure, Software Development & Engineering · Depth: Advanced, medium

Summary

Together has released a major update to its inference platform, offering users complete control over performance, cost, and quality for open-weight AI models without requiring a custom infrastructure stack. This enhanced platform enables models to go live in minutes, supporting production-grade deployments with features like multiple deployments behind a single stable endpoint, safe changes via canary, blue-green, and rolling updates with auto-rollback, and real-traffic testing through A/B and shadow testing. The service also provides flexible autoscaling across regions and offers optimized deployment profiles, leading to approximately 4x faster warm starts. Additionally, Together is launching a closed beta for custom training, including full-weight and LoRA reinforcement learning and supervised fine-tuning, allowing direct deployment of checkpoints to production. The platform aims to simplify the entire model lifecycle, from training to continuous improvement.

Key takeaway

For MLOps Engineers managing open-weight AI model inference, Together's updated platform simplifies complex deployments and continuous iteration. You can achieve production-grade stability and performance in minutes, leveraging features like safe rollouts with auto-rollback and real-traffic testing. This allows you to maintain complete control over your models' performance, quality, and cost, significantly reducing the operational overhead of building and maintaining custom inference stacks. Consider exploring the custom training beta to streamline your model development lifecycle.

Key insights

The platform provides comprehensive control and simplified deployment for open-weight AI inference, integrating training to production.

Principles

Method

The platform supports deploying models from various sources, selecting optimized profiles, scaling based on diverse metrics, and safely rolling out changes with A/B or shadow testing.

In practice

Topics

Best for: CTO, VP of Engineering/Data, Director of AI/ML, MLOps Engineer, Machine Learning Engineer, AI Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Together AI | The AI Native Cloud - Together.ai.