Benchmark ResultNew

OpenAI GPT-6 Astra evaluated on prompt injection resilience

According to The Decoder, OpenAI's GPT-6 Astra reduces hallucinations and blocks 99.99 percent of direct prompt injections, but remains vulnerable to hidden prompt injection attacks embedded within documents in 8.5 percent of evaluated scenarios. Claude Opus 5 demonstrated a lower failure rate of 4.8 percent under similar document-based injection testing. Claims are as reported; this summary makes no determination about accuracy or significance.

First detected
Sep 5, 2026
Last updated
Sep 5, 2026

Moderate confidence

Based on a single independent report.

Limited corroboration

1 reporting source

What does this mean?

Corroboration measures how many genuinely independent sources support the event. Confidence measures how reliable the available evidence appears.

Stable

No recent reporting has materially changed the known facts.

Follow this development to see meaningful updates as new evidence emerges.

Save keeps this for later. Follow tracks meaningful changes as new evidence emerges — it shapes your Following Feed, alerts, and digest eligibility, and doesn't promise an instant notification.

Why it matters

Residual failure rates against indirect prompt injections create operational security exposures for autonomous AI agents handling untrusted external data and documents. Enterprise deployments operating multi-step or autonomous workflows cannot rely solely on native frontier model defenses to neutralize document-borne payload risks.

Coverage

How this developed

  1. Sep 5, 2026

    1. Development detected

  2. Sep 4, 2026

    1. New reporting added