evald.ai Entities

OpenAI Evaluation Filter

Inside our approach to the Model Spec

Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.

Entities: OpenAI

Mitchell Bryson AI Reliability Articles

The land grab has gone financial - Mitchell Bryson

OpenAI's 17.5% guaranteed-return PE pitch, its 450,000 sq ft campus lease, and the Helion fusion deal all point to the same shift: the AI race is no longer about who has the best model — it's about who can lock in distribution, real estate, and energy...

Topics: Benchmarks

Entities: AnthropicBenchmarksOpenAI

OpenAI Evaluation Filter

Introducing EVMbench

OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-severity smart contract vulnerabilities.

Topics: Benchmarks

Entities: BenchmarksOpenAI

OpenAI Evaluation Filter

Evaluating chain-of-thought monitorability

OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations across 24 environments. Our findings show that monitoring a model’s internal reasoning is far more effective than monitoring outputs alone,...

Entities: OpenAI

OpenAI Evaluation Filter

Updating our Model Spec with teen protections

OpenAI is updating its Model Spec with new Under-18 Principles that define how ChatGPT should support teens with safe, age-appropriate guidance grounded in developmental science. The update strengthens guardrails, clarifies expected model behavior in...

Topics: Safety Evals

Entities: Safety EvalsOpenAIChatGPT