LLMs break down in funny ways when told the Jacobian Conjecture counterargument

· Source: Max Woolf's Blog · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Mathematics & Computational Sciences, Emerging Technologies & Innovation · Depth: Intermediate, medium

Summary

An Anthropic researcher, Levent Alpöge, recently identified a counterargument to the 80-year-old Jacobian Conjecture using Claude Fable 5. This discovery, empirically validated, created a logical paradox for large language models (LLMs) because their knowledge bases are locked prior to July 19th, 2026, when the conjecture was still unsolved. The author tested 14 different LLMs, including GPT-5.6 Sol, Claude Opus 4.8, and Gemini 3.5 Flash, with a specific mathematical query presenting the counterargument. Seven models, such as Gemini 3.5 Flash and Qwen3.7 Max, confirmed and proved the counterargument, while five argued against it. MiniMax M3 exceeded its response length, and Claude Opus 4.8 believed the counterargument already existed. When the prompt implied a "cat" found the proof, models exhibited "snark" and humor, with some even suggesting the cat deserved a Fields Medal.

Key takeaway

For AI Scientists evaluating LLM reliability, you should rigorously test models with novel, verified information that challenges their training data. This reveals how models handle logical paradoxes and factual inconsistencies arising from knowledge cutoffs. Consider using diverse prompts, like the "cat" scenario, to assess an LLM's ability to maintain a consistent persona and verify complex mathematical claims under unexpected conditions.

Key insights

LLMs face a logical paradox when presented with new, verified mathematical truths contradicting their training data.

Principles

Method

The article describes a method of testing LLMs by providing a known mathematical counterargument and observing their verification capabilities and reactions to information contradicting their training data.

In practice

Topics

Code references

Best for: Research Scientist, AI Engineer, AI Scientist, Machine Learning Engineer, Prompt Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Max Woolf's Blog.