Google adopts Werewolf and Poker in AI benchmark 'Game Arena' - GIGAZINE
Google adopts Werewolf and Poker in AI benchmark 'Game Arena' GIGAZINE
Topics: Benchmarks
Entities: BenchmarksGoogle
Company
Google adopts Werewolf and Poker in AI benchmark 'Game Arena' GIGAZINE
Topics: Benchmarks
Entities: BenchmarksGoogle
GPT-5.2 lands to top Google's Gemini 3 in the AI benchmark game just four weeks after GPT-5.1 the-decoder.com
Topics: Benchmarks
Entities: BenchmarksGeminiGoogle
Google DeepMind and the UK AI Security Institute (AISI) strengthen collaboration through a new research partnership, focusing on critical safety research areas like monitoring AI reasoning and evalua…
Topics: Safety Evals
Entities: Safety EvalsGoogleGoogle DeepMind
Revolutionary Google Gemini 3 Shatters Records with Unprecedented AI Benchmark Scores and Game-Changing Coding App CryptoRank
Topics: Benchmarks
Entities: BenchmarksGeminiGoogle
Introducing Metrax: performant, efficient, and robust model evaluation metrics in JAX blog.google
Topics: LLM Evaluation
Entities: LLM EvaluationGoogle
Today, we’re publishing the third iteration of our Frontier Safety Framework (FSF) — our most comprehensive approach yet to identifying and mitigating severe risks from advanced AI models. This updat…
Entities: GoogleGoogle DeepMind
Google Stax Aims to Make AI Model Evaluation Accessible for Developers infoq.com
Topics: LLM Evaluation
Entities: LLM EvaluationGoogle
Explore Stax, an experimental developer tool that streamlines LLM evaluation with human labelling and scalable LLM-as-a-judge auto-raters for data driven decisions.
Topics: LLM EvaluationTesting Tools
Entities: Testing ToolsLLM EvaluationGoogle
Vincent Cheng and Thomas Kwa replicate a Google DeepMind paper on chain-of-thought monitoring, showing evidence that monitoring works on other companies' models.
Entities: ClaudeGeminiGoogleGoogle DeepMind