evald.ai Sources

OpenAI Evaluation Filter

New funding to build towards AGI

Today we’re announcing new funding—$40B at a $300B post-money valuation, which enables us to push the frontiers of AI research even further, scale our compute infrastructure, and deliver increasingly powerful tools for the 500 million people who use ChatGPT...

Entities: ChatGPT

OpenAI Evaluation Filter

Deep research System Card

This report outlines the safety work carried out prior to releasing deep research including external red teaming, frontier risk evaluations according to our Preparedness Framework, and an overview of the mitigations we built in to address key risk areas.

Topics: Safety Evals

Entities: Safety Evals

OpenAI Evaluation Filter

OpenAI o3-mini System Card

This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external red teaming, and Preparedness Framework evaluations.

Topics: Safety Evals

Entities: Safety EvalsOpenAI

OpenAI Evaluation Filter

Operator System Card

Drawing from OpenAI’s established safety frameworks, this document highlights our multi-layered approach, including model and product mitigations we’ve implemented to protect against prompt engineering and jailbreaks, protect privacy and security, as well...

Entities: OpenAI

OpenAI Evaluation Filter

OpenAI o1 System Card

This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations according to our Preparedness Framework.

Topics: Safety Evals

Entities: Safety EvalsOpenAI

OpenAI Evaluation Filter

Introducing SimpleQA

A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.

Topics: Benchmarks

Entities: Benchmarks