Qwen3 7 Max Adds Speed And Power
Summary
Alibaba has updated its flagship large language model, Qwen3.7-Max, positioning it for text-only tasks such as coding and scientific discovery. This model, with closed weights, accepts up to 1 million input tokens and generates up to 64,000 output tokens at 208.3 tokens per second. Key features include reasoning, tool use, prompt caching, and native compatibility with OpenAI and Anthropic API specifications. Qwen3.7-Max ranks seventh on the Artificial Analysis Intelligence Index with a reasoning score of 56.6 and sixth on AA-Omniscience (14), while achieving the third-fastest output speed. It is available free via Qwen Chat and through Alibaba Cloud Model Studio API at \$2.50/\$0.25/\$7.50 per million input/cached/output tokens. Alibaba also claims strong agentic capabilities, demonstrated by an internal test where it optimized an attention kernel 10 times faster. This release continues Alibaba's trend of shifting its top-tier models to closed weights for revenue generation.
Key takeaway
For Machine Learning Engineers evaluating LLMs for agentic applications or high-throughput text processing, Qwen3.7-Max presents a compelling option. Its third-fastest speed and demonstrated agentic capabilities, like 10x kernel optimization, make it suitable for demanding tasks. You should consider its API via Alibaba Cloud Model Studio for cost-effective long-running workflows, acknowledging its closed-source nature and the undisclosed training details.
Key insights
Alibaba's Qwen3.7-Max demonstrates high performance and speed, particularly for agentic tasks, despite its closed-source nature.
Principles
- Agent training benefits from decoupled components.
- Abstention can improve output correctness.
- Closed models can drive revenue.
Method
Alibaba's reinforcement learning approach for Qwen3.7-Max separates task, agentic harness, and verifier components, training on diverse combinations to prevent setup-specific learning.
In practice
- Utilize Qwen3.7-Max for long-running agentic workflows.
- Integrate with OpenAI/Anthropic APIs for compatibility.
- Consider Qwen Chat for free access.
Topics
- Qwen3.7-Max
- Large Language Models
- Agentic AI
- Reinforcement Learning
- API Integration
- Model Performance Benchmarks
Best for: AI Engineer, NLP Engineer, CTO, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by The Batch | DeepLearning.AI | AI News & Insights - www.deeplearning.ai.