The Open Arabic LLM Leaderboard 2
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
Topic feed
AI benchmarks, leaderboards, and comparative model testing.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
What Makes a Good AI Benchmark? Stanford HAI
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: BenchmarksLLM Evaluation
Entities: BenchmarksLLM Evaluation
We’re releasing RE-Bench, a new benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. We also share data from 71 human expert attempts and results for Anthropic’s Claude 3.5 Sonnet and OpenAI’s...
Topics: Benchmarks
Entities: AnthropicClaudeBenchmarksOpenAI
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Topics: Benchmarks
Entities: Benchmarks
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Topics: Benchmarks
Entities: Benchmarks