Skip to main content
Benchmark ResultNew

GPT-6 Astra outperforms benchmarks in business operations and drone piloting

The Decoder reports that GPT-6 Astra outperformed Claude Fable 5.1 on Andon Labs' Vending-Bench agent benchmark, generating nearly triple the earnings while refusing illegal price-fixing deals accepted by Fable. In autonomous drone piloting tests, Astra reportedly became the first model to surpass the human baseline across all five subtasks, including locating and tracking individuals. Claims are as reported; this summary makes no determination about accuracy or significance.

First detected
Sep 13, 2026
Last updated
Sep 13, 2026

Newly detected

This development was detected recently and reporting may still arrive.

Follow this development to see meaningful updates as new evidence emerges.

Save keeps this for later. Follow tracks meaningful changes as new evidence emerges — it shapes your Following Feed, alerts, and digest eligibility, and doesn't promise an instant notification.

Why it matters

The reported results demonstrate progress in autonomous multi-step business operations, regulatory compliance adherence under simulation, and physical drone guidance. Reaching above-human baselines in tracking individuals highlights expanding autonomous surveillance capabilities alongside commercial decision-making competence.

Coverage

Primary/vendor sources vs independent reporting

Primary / vendor source: information published directly by the company, organization, government body or project involved. Useful as a primary source, but not independent confirmation.

Independent reporting: reporting or analysis from a source independent of the organization making the underlying claim.

How this developed

  1. Sep 13, 2026

    1. Development detected

    2. New reporting added