ModelsResearch

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

Source: The Decoder · Maximilian Schreiner

Intel Summary

OpenAI's GPT-6 Astra has generated conflicting benchmark results across evaluation platforms. Epoch AI ranked the model in the lead with 169 points, whereas Artificial Analysis evaluated it on par with its predecessor and behind Claude Fable 5.1. However, on ARC-AGI-3, Astra operated more efficiently than the average human, leading ARC Prize lead François Chollet to advance his AGI timeline after observing progress moving twice as fast as projected.

Why It Matters

Conflicting evaluation metrics highlight growing divergence in how independent tracking organizations assess frontier AI performance across standard workloads versus abstract reasoning. Demonstrating superior efficiency to humans on ARC-AGI-3 accelerates timelines for complex problem-solving capabilities, even as conventional benchmark improvements appear uneven across competing frontier models.

Part of an ongoing development

Independent reporting

OpenAI GPT-6 Astra evaluated across frontier benchmarks

OpenAI's GPT-6 Astra has generated conflicting benchmark results across evaluation platforms. Epoch AI ranked the model in the lead with 169 points, whereas Artificial Analysis evaluated it on par with its predecessor and behind Claude Fable 5.1. Claims are as reported; this summary makes no determination about accuracy or significance.

Confidence
Moderate confidence
Corroboration
Limited corroboration

Organizations & Entities