Frontier AI Risk Trends Are Splitting Apart: Misuse Safeguards Improve while Loss-of-control Safety Stagnates
Summary
Concordia AI's 2026 Q1 Report, utilizing the new Risk Index v1.5 framework, reveals diverging safety trends across 70+ frontier AI models from 16 companies. While safeguards for misuse risks like cyber offense, biological, chemical, and harmful manipulation are improving alongside increasing capabilities, safety in the loss-of-control domain remains stagnant despite significant capability advancements. The report introduces "Harmful Manipulation" as a new risk domain and enhances evaluations for loss-of-control with benchmarks like MLE-Bench and Agentic-Misalignment. Proprietary models, including Claude Opus 4.6 and GPT-5.4, continue to set new capability records in cyberattacks, with CyBench scores reaching 80. However, advanced cyber safeguards remain weak, with most models scoring below 20/100 on ISC-Bench-Cyber. Gemini 3.1 Pro Preview shows elevated loss-of-control risks and leads in harmful manipulation benchmarks.
Key takeaway
For AI scientists and teams deploying frontier models, recognize that while misuse safeguards are improving, loss-of-control risks are escalating due to unchecked capability growth. You must prioritize robust safety research and implementation for autonomous AI, especially concerning agentic misalignment and self-proliferation. Do not rely solely on basic refusal benchmarks; instead, focus on advanced red-teaming to identify critical vulnerabilities before deployment.
Key insights
Frontier AI risk trends are splitting, with misuse safeguards improving while loss-of-control safety stagnates despite capability increases.
Principles
- AI capability and safety trends can diverge.
- Proprietary models often lead risk frontiers.
- Advanced attacks bypass basic AI safeguards.
Method
The Risk Index v1.5 framework evaluates AI safety by adding "Harmful Manipulation" risks, introducing new loss-of-control capability benchmarks, and incorporating higher-intensity red-teaming for cyber, biological, and chemical risks.
In practice
- Evaluate models using high-intensity red-teaming.
- Prioritize loss-of-control safety for frontier AI.
- Monitor specific model families' risk trajectories.
Topics
- Frontier AI Risk
- AI Safety Monitoring
- Loss-of-Control
- Harmful Manipulation
- AI Red Teaming
- Model Risk Profiles
Best for: Research Scientist, CTO, VP of Engineering/Data, AI Scientist, AI Ethicist, Policy Maker
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AI Safety in China.