Piloting the world's first double-blind AI evaluations
Building trust in proprietary model benchmarks using cryptographically secure environments
Google DeepMind ยท William Isaac, Sol Messing and Kristian Lum
Topics: Benchmarks
Entities: Benchmarks
Building trust in proprietary model benchmarks using cryptographically secure environments
Google DeepMind ยท William Isaac, Sol Messing and Kristian Lum
Topics: Benchmarks
Entities: Benchmarks