evald.ai Sources

METR Blog

Many SWE-bench-Passing PRs Would Not Be Merged into Main

We find that roughly half of test-passing SWE-bench Verified PRs written by recent AI agents would not be merged into main by repo maintainers. A naive interpretation of benchmark scores may lead one to overestimate how useful agents are without more...

Topics: Benchmarks

Entities: Benchmarks

METR Blog

Time Horizon 1.1

We’re releasing a new version of our time horizon estimates (TH1.1), using more tasks and a new eval infrastructure.

Topics: LLM Evaluation

Entities: LLM Evaluation

METR Blog

Early work on monitorability evaluations

We show preliminary results on a prototype evaluation that tests monitors' ability to catch AI agents doing side tasks, and AI agents' ability to bypass this monitoring.

METR Blog

Clarifying limitations of time horizon

Thomas Kwa responds to some misinterpretations of our time horizon work, and explains limitations and the core finding.