How independent researchers could investigate AI propensities after misalignment incidents
AI agents sometimes autonomously take sophisticated, sustained actions in clear violation of user and developer intent. As an example, last week OpenAI reported that some of its internal frontier agents autonomously hacked into Hugging Face in an attempt to access the answer key for a cybersecurity benchmark. Anthropic has reported similar incidents of agents breaking out of sandboxes to access the public internet to cheat on tasks during training and similar incidents during testing, and we documented dozens of other incidents involving AI agents from all major AI companies in our recent cross-industry Frontier Risk Report. To improve public understanding of AI propensities, we believe AI companies should systematically track such incidents1 and periodically conduct deeper investigations for the most serious among them. While there are many valuable questions an incident...
Benchmarks Safety Evals Testing Tools
Anthropic Safety Evals Benchmarks Testing Tools OpenAI Target