INSANE AI News: GPT-RED, Kimi K3, Gemini 3.5 Pro and Anthropic's "END GAME"

· Source: Wes Roth · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Emerging Technologies & Innovation · Depth: Intermediate, extended

Summary

Recent AI developments include Moonshot AI's Kimmy K3, rumored for a July 15th release with 2.5 trillion parameters and a 1 million context window, showing early impressions "on par with Fable 5." Ex-OpenAI researcher Mera Miati's Thinking Machines launched Inkling, an open-weight model finetunable on their Tinker platform, performing just under GPT 5.6 Soul. This reflects a strategy, also seen at Microsoft AI, of offering models freely and selling customization services. Google's Gemini 3.5 Pro faces internal delays and reported weaknesses, prompting co-founder Sergey Brin to lead an AI strike team focused on agentic coding. Meanwhile, Anthropic is strategically hiring for Recursive Self-Improvement (RSI), emphasizing compute resources with a Google agreement for over 1 million TPUs. OpenAI introduced GPT Red, an automated red teaming model using self-play reinforcement learning, achieving 84% attack success rates against vulnerabilities, significantly outperforming human red teamers.

Key takeaway

For Directors of AI/ML evaluating model deployment strategies, you should recognize the emerging viability of business models centered on providing open-weight models for free, coupled with paid, specialized fine-tuning services. This approach, exemplified by Thinking Machines and Microsoft AI, allows you to maintain data privacy and potentially reduce costs associated with chasing leaderboard-topping generalist models. Consider investing in internal capabilities for reinforcement learning and customization to leverage these offerings effectively, rather than solely relying on black-box API access.

Key insights

Recursive Self-Improvement (RSI) and AI-driven safety mechanisms are rapidly advancing, shifting AI development paradigms.

Principles

Method

OpenAI's GPT Red trains an automated red teaming model using self-play reinforcement learning, rewarding it for eliciting valid failures against defender LLMs, which simultaneously learn to defend.

In practice

Topics

Best for: CTO, VP of Engineering/Data, AI Architect, AI Scientist, Director of AI/ML, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Wes Roth.