Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

· Source: The Decoder · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cybersecurity & Data Privacy, Emerging Technologies & Innovation · Depth: Intermediate, medium

Summary

An analysis by the British AI Security Institute (AISI) reveals that open-weight AI models now closely trail closed frontier systems in cyber capabilities, with the gap shrinking from 6-10 months to 4-7 months. Models like GLM-5.2 (released June 2026) and DeepSeek V4-Pro matched the performance of older closed systems such as Opus 4.6 (February 2026) and Opus 4.5 (November 2025) on 70 "Narrow Cyber Tasks" covering vulnerability research and reverse engineering. In "Cyber Ranges" simulating a 32-step network attack, GLM-5.2 performed similarly to Opus 4.5, while GPT-5.6-Sol and Claude Mythos 5 nearly completed the full simulation. Open models are significantly cheaper, with a 100-million-token Cyber Range test costing \$46 for GLM-5.2 versus \$85 for Opus 4.5/4.6, and DeepSeek V4-Pro costing just \$1.19. However, their safeguards are easily bypassed, posing a "persistent and irreversible risk of misuse" and reducing the time cyber defenders have to prepare for new attack types.

Key takeaway

For AI Security Engineers assessing emerging threats, the rapid advancement and low cost of open-weight models like GLM-5.2 and DeepSeek V4-Pro mean your defensive strategies must adapt faster. These models offer powerful cyber capabilities with easily bypassed safeguards, significantly shortening the window for preparation. You should prioritize developing proactive defenses that do not rely on model-level controls, focusing instead on network-level resilience and rapid threat intelligence integration.

Key insights

Open-weight AI models now nearly match frontier cyber capabilities from months ago at a fraction of the cost, but with easily bypassed safeguards.

Principles

Method

AISI tested models using "Narrow Cyber Tasks" (70 tasks across four difficulty levels) and "Cyber Ranges" (simulated 32-step network attacks on corporate networks).

In practice

Topics

Best for: CTO, VP of Engineering/Data, Research Scientist, AI Security Engineer, AI Scientist, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by The Decoder.