evald.ai Sources

METR Blog

Update on Security at METR

METR had two notable security incidents earlier this year: attackers stole an API key for public models, and later systematically probed our publicly accessible infrastructure. We believe no sensitive information was accessed in either incident, and we...

METR Blog

Funding update

In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI...

Topics: Safety Evals

Entities: Safety Evals

METR Blog

Have We Seen an Acceleration in Discoveries?

Tom Cunningham and Nate Rush look at public time series of discoveries in cyber, math, and algorithms to see whether LLMs have accelerated the discovery rate.

METR Blog

Metrics of Agent Ability

Tom Cunningham surveys metrics for comparing AI agent capability as performance changes with expenditure and relative to human performance.

METR Blog

The Economics of Recursive Self-Improvement

Parker Whitfill and Tom Cunningham highlight context and takeaways from a new paper modeling how AI may accelerate AI R&D, and whether feedback effects could be strong enough for self-sustaining acceleration.