Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
OpenAI's GPT-6 Astra has generated conflicting benchmark results across evaluation platforms. Epoch AI ranked the model in the lead with 169 points, whereas Artificial Analysis evaluated it on par with its predecessor and behind Claude Fable 5.1. However, on ARC-AGI-3, Astra operated more efficiently than the average human, leading ARC Prize lead François Chollet to advance his AGI timeline after observing progress moving twice as fast as projected.