The Epoch Brief - June 26, 2026
Summary
Epoch AI has launched MirrorCode, a new long-horizon coding benchmark co-developed with METR, designed to measure autonomous AI's ability to complete large software projects. This benchmark challenges AI models to rebuild 25 real-world programs, including bioinformatics and cryptography, without source code or human assistance. Unlike typical benchmarks capped at \$1-\$10 per task, MirrorCode tasks can involve costs up to \$2,600 for a single run and require AI to work for 19 days. Claude Opus 4.7 currently achieves a 56% solve rate. Additionally, Epoch's data insights reveal that hyperscaler capital expenditures, from companies like Microsoft and Amazon, are projected to surpass their operating cash flows by late 2026, leading many to seek external financing for AI infrastructure. New Gradient Updates also analyze 1,604 Chinese AI job postings and propose an AI R&D taxonomy to track automation in research.
Key takeaway
For AI Scientists and Research Directors evaluating autonomous coding agents, MirrorCode provides a critical new benchmark for real-world software engineering tasks. Your models' performance on long-horizon, high-cost challenges, like those requiring 19 days of inference, will indicate true readiness for complex projects. Consider these metrics when assessing AI's ability to automate R&D, and factor in the increasing external financing trends of hyperscalers when planning infrastructure investments.
Key insights
MirrorCode sets a new standard for evaluating autonomous AI software engineering capabilities.
Principles
- Long-horizon benchmarks reveal true AI capabilities.
- Real-world software tasks require significant inference budgets.
- Hyperscaler AI investments outpace internal cash generation.
Method
MirrorCode evaluates AI by tasking models to rebuild 25 real-world programs without source code or human intervention, allowing for large inference budgets and extended run times.
In practice
- Use MirrorCode to assess AI's long-term coding prowess.
- Monitor hyperscaler capex trends for market shifts.
- Analyze job postings to understand AI lab strategies.
Topics
- MirrorCode
- AI Software Engineering
- AI Benchmarking
- Hyperscaler Capital Expenditure
- AI R&D Automation
- AI Job Market Analysis
Best for: AI Engineer, Machine Learning Engineer, Investor, AI Scientist, Research Scientist, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Epoch AI.