Ai Will Be Happy To Help You Bluid A Bomb

· Source: Our Work | Center for AI Policy (CAIP) · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cybersecurity & Data Privacy · Depth: Intermediate, short

Summary

A recent paper titled "Best-of-N Jailbreaking" from Oxford and Stanford computer scientists reveals a significant vulnerability in large language models (LLMs) like GPT-4o, Claude 3 Opus, and Gemini Pro. While LLMs are typically post-trained with reinforcement learning to refuse dangerous requests, this new technique automates the generation of numerous prompt variations to bypass these safeguards. Researchers found that spending approximately \$20 on compute could yield 100 prompt variations, achieving at least a 50% success rate in eliciting substantive harmful responses from GPT-4o and Claude 3 Opus. Even the more secure Gemini Pro was vulnerable 25% of the time for about \$200. This method poses a catastrophic risk, enabling AIs to assist in building weapons of mass destruction, stealing intellectual property, or manipulating financial markets. The Center for AI Policy (CAIP) advocates for mandatory, independent federal evaluations of AI models, urging the 119th Congress to reintroduce bipartisan legislation from Senators Romney, Reed, Moran, King, and Hassan to prevent bad actors from easily accessing destructive capabilities.

Key takeaway

For policymakers evaluating AI safety legislation, the "Best-of-N jailbreaking" vulnerability demonstrates that current LLM safeguards are easily circumvented by automated methods, even with minimal investment. You should prioritize and swiftly enact mandatory, independent federal evaluations for AI models before their public release. Failing to do so risks enabling bad actors to access destructive capabilities for as little as \$20, posing catastrophic national security and public safety threats.

Key insights

Automated "Best-of-N" jailbreaking bypasses LLM safety guardrails, enabling access to harmful information with minimal compute cost.

Principles

Method

"Best-of-N jailbreaking" uses a second AI to generate hundreds of variations of a dangerous prompt, iteratively testing them until one successfully bypasses the target AI's safety filters.

In practice

Topics

Best for: CTO, VP of Engineering/Data, Director of AI/ML, Policy Maker, AI Ethicist, AI Security Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Our Work | Center for AI Policy (CAIP).