Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6
External review from METR of Anthropic's Sabotage Risk Report for Claude Opus 4.6
Topics: Safety Evals
Entities: AnthropicSafety EvalsClaudeClaude Opus
Concept
External review from METR of Anthropic's Sabotage Risk Report for Claude Opus 4.6
Topics: Safety Evals
Entities: AnthropicSafety EvalsClaudeClaude Opus
OpenAI amends its Pentagon deal after Altman admits it looked 'opportunistic and sloppy', while Claude surges to number one on the App Store and hundreds of employees publicly back Anthropic's stance.
Topics: Safety Evals
Entities: AnthropicSafety EvalsClaudeOpenAI
Defense Secretary Pete Hegseth gives Anthropic until Friday to provide military access to Claude or face being declared a supply chain risk or forced compliance under the Defense Production Act.
Topics: Safety Evals
Entities: AnthropicSafety EvalsClaude
Luca Righetti shares takeaways on the role of randomized controlled trials in AI safety testing.
Topics: Safety EvalsTesting Tools
Entities: Safety EvalsTesting Tools
Miles Kodama and Michael Chen summarize key provisions from California's SB 53, the EU Code of Practice, and New York's RAISE Act covering frontier AI developers.
Topics: Safety Evals
Entities: Safety Evals
OpenAI is updating its Model Spec with new Under-18 Principles that define how ChatGPT should support teens with safe, age-appropriate guidance grounded in developmental science. The update strengthens guardrails, clarifies expected model behavior in...
Topics: Safety Evals
Entities: Safety EvalsOpenAIChatGPT
Announcing Gemma Scope 2, a comprehensive, open suite of interpretability tools for the entire Gemma 3 family to accelerate AI safety research.
Topics: Safety Evals
Entities: Safety Evals
Google DeepMind and the UK AI Security Institute (AISI) strengthen collaboration through a new research partnership, focusing on critical safety research areas like monitoring AI reasoning and evalua…
Topics: Safety Evals
Entities: Safety EvalsGoogleGoogle DeepMind
Shared components of AI lab commitments to evaluate and mitigate severe risks.
Topics: Safety Evals
Entities: Safety Evals
A practical model for growing AI agent autonomy - levels, controls, and KPIs - grounded in risk frameworks (NIST AI RMF, ISO/IEC 42001) and aligned with EU AI Act oversight.
Topics: Safety Evals
Entities: Safety Evals
External review from METR of Anthropic's Summer 2025 Sabotage Risk Report
Topics: Safety Evals
Entities: AnthropicSafety Evals
Details on external recommendations from METR for gpt-oss Preparedness experiments and follow-up from OpenAI.
Topics: Safety Evals
Entities: Safety EvalsOpenAI