Common Elements of Frontier AI Safety Policies (December 2025 Update)
Shared components of AI lab commitments to evaluate and mitigate severe risks.
Topics: Safety Evals
Entities: Safety Evals
Shared components of AI lab commitments to evaluate and mitigate severe risks.
Topics: Safety Evals
Entities: Safety Evals
We evaluate whether GPT-5.1-Codex-Max poses significant catastrophic risks via AI self-improvement, rogue replication, or sabotage of AI labs. We conclude that this seems unlikely.
Entities: OpenAI
External review from METR of Anthropic's Summer 2025 Sabotage Risk Report
Topics: Safety Evals
Entities: AnthropicSafety Evals
Details on external recommendations from METR for gpt-oss Preparedness experiments and follow-up from OpenAI.
Topics: Safety Evals
Entities: Safety EvalsOpenAI
MALT (Manually-reviewed Agentic Labeled Transcripts) is a dataset of natural and prompted examples of behaviors that threaten evaluation integrity (like generalized reward hacking or sandbagging).
Topics: LLM Evaluation
Entities: LLM Evaluation
Vincent Cheng, Thomas Kwa, and Neev Parikh share research on how AI agents can hide secondary task-solving from monitors, finding that harder tasks are more detectable and small models can learn to evade larger monitors.
Vincent Cheng and Thomas Kwa replicate a Google DeepMind paper on chain-of-thought monitoring, showing evidence that monitoring works on other companies' models.
Entities: ClaudeGeminiGoogleGoogle DeepMind
AI agents are improving rapidly at autonomous software development and machine learning tasks, and, if recent trends hold, may match human researchers at challenging months-long research projects in under a decade. Some economic models predict that...
Many AI benchmarks use algorithmic scoring to evaluate how well AI systems perform on some set of tasks. However, AI systems often produce code that scores well but isn't production-ready due to issues with test coverage, formatting, and code quality. This...
Topics: BenchmarksLLM Evaluation
Entities: BenchmarksLLM Evaluation
How we think about tradeoffs when communicating surprising or nuanced findings.
Recent work from Anthropic and others claims that LLMs' chains of thoughts can be “unfaithful”. These papers make an important point: you can't take everything in the CoT at face value. As a result, people often use these results to conclude the CoT is...
Entities: Anthropic
We evaluate whether GPT-5 poses significant catastrophic risks via AI self-improvement, rogue replication, or sabotage of AI labs. We conclude that this seems unlikely. However, capability trends continue rapidly, and models display increasing eval awareness.
Topics: LLM Evaluation
Entities: LLM EvaluationOpenAI