BrowseComp: a benchmark for browsing agents
BrowseComp: a benchmark for browsing agents.
Topics: Benchmarks
Entities: Benchmarks
BrowseComp: a benchmark for browsing agents.
Topics: Benchmarks
Entities: Benchmarks
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
Topics: Benchmarks
Entities: Benchmarks
Today we’re announcing new funding—$40B at a $300B post-money valuation, which enables us to push the frontiers of AI research even further, scale our compute infrastructure, and deliver increasingly powerful tools for the 500 million people who use ChatGPT...
Entities: ChatGPT
This report outlines the safety work carried out prior to releasing deep research including external red teaming, frontier risk evaluations according to our Preparedness Framework, and an overview of the mitigations we built in to address key risk areas.
Topics: Safety Evals
Entities: Safety Evals
Can frontier LLMs earn $1 million from real-world freelance software engineering?
Topics: Benchmarks
Entities: Benchmarks
This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external red teaming, and Preparedness Framework evaluations.
Topics: Safety Evals
Entities: Safety EvalsOpenAI
Drawing from OpenAI’s established safety frameworks, this document highlights our multi-layered approach, including model and product mitigations we’ve implemented to protect against prompt engineering and jailbreaks, protect privacy and security, as well...
Entities: OpenAI
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations according to our Preparedness Framework.
Topics: Safety Evals
Entities: Safety EvalsOpenAI
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Topics: Benchmarks
Entities: Benchmarks
We’ve simplified, stabilized, and scaled continuous-time consistency models, achieving comparable sample quality to leading diffusion models, while using only two sampling steps.
Topics: LLM Evaluation
Entities: LLM Evaluation
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Topics: Benchmarks
Entities: Benchmarks
OpenAI and Los Alamos National Laboratory are working to develop safety evaluations to assess and measure biological capabilities and risks associated with frontier models.
Entities: OpenAI