ModelsResearchSecurity

AI benchmarks have a trust problem and Google wants to fix it

Source: The Decoder · Maximilian Schreiner

Intel Summary

Google DeepMind has launched a pilot project with the Singapore AI Safety Institute to conduct double-blind evaluations of frontier AI models. Using Google's Confidential Space cryptographic environment, the framework prevents Google from accessing evaluation benchmark datasets while keeping model weights protected from external evaluators. Tested on Gemini Flash Lite, the initiative aims to establish tamper-proof, contamination-resistant testing standards for AI model safety and capability assessments.

Why It Matters

AI benchmarks suffer from severe dataset contamination and gaming risks, complicating independent safety verification and enterprise model selection. If successful, cryptographically isolated evaluation protocols could become an industry standard for regulatory audits and commercial validation, enabling model developers and external safety bodies to verify capabilities without exposing proprietary intellectual property or test questions.

Part of an ongoing development

Independent reporting

Google DeepMind launched double-blind AI evaluation pilot with Singapore AI Safety Institute

Google DeepMind has launched a pilot project with the Singapore AI Safety Institute to conduct double-blind evaluations of frontier AI models. Using Google's Confidential Space cryptographic environment, the framework prevents Google from accessing evaluation benchmark datasets while keeping model weights protected from external evaluators. Claims are as reported; this summary makes no determination about accuracy or significance.

Confidence
Moderate confidence
Corroboration
Limited corroboration

Organizations & Entities