Benchmark Saturation in 2026: Why Frontier Models Top Humanity's Last Exam & What Still Separates Them
Humanity's Last Exam tops 80% and ARC-AGI-2 scores tripled, yet the production gap between models is wider than ever. Here is why benchmarks saturate, what still separates models, and the eval suite that decides real procurement.