LLMs break down in funny ways when told the Jacobian Conjecture counterargument
Summary
An Anthropic researcher, Levent Alpöge, recently identified a counterargument to the 80-year-old Jacobian Conjecture using Claude Fable 5. This discovery, empirically validated, created a logical paradox for large language models (LLMs) because their knowledge bases are locked prior to July 19th, 2026, when the conjecture was still unsolved. The author tested 14 different LLMs, including GPT-5.6 Sol, Claude Opus 4.8, and Gemini 3.5 Flash, with a specific mathematical query presenting the counterargument. Seven models, such as Gemini 3.5 Flash and Qwen3.7 Max, confirmed and proved the counterargument, while five argued against it. MiniMax M3 exceeded its response length, and Claude Opus 4.8 believed the counterargument already existed. When the prompt implied a "cat" found the proof, models exhibited "snark" and humor, with some even suggesting the cat deserved a Fields Medal.
Key takeaway
For AI Scientists evaluating LLM reliability, you should rigorously test models with novel, verified information that challenges their training data. This reveals how models handle logical paradoxes and factual inconsistencies arising from knowledge cutoffs. Consider using diverse prompts, like the "cat" scenario, to assess an LLM's ability to maintain a consistent persona and verify complex mathematical claims under unexpected conditions.
Key insights
LLMs face a logical paradox when presented with new, verified mathematical truths contradicting their training data.
Principles
- LLM knowledge cutoffs create factual inconsistencies.
- LLMs can verify complex mathematical proofs.
- Contextual framing influences LLM persona.
Method
The article describes a method of testing LLMs by providing a known mathematical counterargument and observing their verification capabilities and reactions to information contradicting their training data.
In practice
- Test LLMs with novel, verified information.
- Observe LLM behavior under logical paradox.
- Evaluate LLM persona consistency.
Topics
- Jacobian Conjecture
- Large Language Models
- LLM Evaluation
- Mathematical Proof Verification
- Knowledge Cutoffs
- AI Paradoxes
Code references
Best for: Research Scientist, AI Engineer, AI Scientist, Machine Learning Engineer, Prompt Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Max Woolf's Blog.