evald.ai Entities

Mitchell Bryson AI Reliability Articles

The leak that repriced cybersecurity - Mitchell Bryson

Anthropic's accidental Mythos reveal crashed cybersecurity stocks — but the market was catching up to a reality that was already here. On the same day, CISA warned of active exploitation of AI agent frameworks, researchers disclosed basic vulnerabilities in...

Entities: AnthropicMythos

Mitchell Bryson AI Reliability Articles

The land grab has gone financial - Mitchell Bryson

OpenAI's 17.5% guaranteed-return PE pitch, its 450,000 sq ft campus lease, and the Helion fusion deal all point to the same shift: the AI race is no longer about who has the best model — it's about who can lock in distribution, real estate, and energy...

Topics: Benchmarks

Entities: AnthropicBenchmarksOpenAI

Hacker News LLM Evaluation

A Synthesis of LLM Evaluation | Arnab Roy

I have been reading a ton about LLM evaluation practices over the past few weeks from Anthropic’s engineering blog, Hamel Husain’s practitioner-focused guides, the Evals for AI Engineers book by Shreya Shankar and Hamel Husain, and several eval framework...

Topics: LLM Evaluation

Entities: AnthropicLLM Evaluation