Benchmarks for frontier, agentic, and safety capabilities
Including benchmarks on agentic coding, frontier reasoning, and safety alignment.
From leading AI labs including OpenAI, Anthropic, Google, Meta, and open-source contributors.
DrugDiscoveryBench: 82 expert-curated tasks evaluating how reliably frontier coding agents perform the computational work of early-stage drug discovery.
GPT-5.5 (mini-SWE-agent) xhigh
51.60±4.30
Claude Sonnet 5 (mini-swe-agent) max
50.04±0.70
Gemini 3.5 Flash (Gemini CLI) high
50.00±2.40
Evaluating an agent's ability to restructure code while preserving behavior.
Fable-5 (Claude Code) xHigh
54.76±6.76
Opus-4.7 (Claude Code)
48.57±6.73
Opus 4.8 (Claude Code)
46.67±6.75
Evaluating an agent’s ability to write production-grade tests
Opus 5 (Claude Code) xHigh
NEW62.22±5.58
Fable-5 (Claude Code) xHigh*
NEW55.60±5.80
Opus 4.8 (Claude Code) xhigh
49.63±5.93
Evaluating deep code comprehension and reasoning
Opus 5 (Claude Code) xHigh
NEW63.17±5.01
Opus 4.8 (Claude Code) xhigh
57.26±4.93
GLM 5.2 (Mini-SWE-Agent)
48.12±5.07
Evaluates whether agents recognize information gaps and ask targeted clarifying questions.
Claude Opus 5
NEW57.00±5.48
Claude Fable 5
56.33±5.50
GLM 5.2
43.67±5.65
Evaluating real-world tool use through the Model Context Protocol (MCP)
Muse Spark 1.1
NEW88.10±1.95
claude-opus-5 (xhigh)
NEW85.80±2.10
gemini-3.5-flash (high)
83.60±2.30
Evaluating long-horizon software engineering tasks in public open source repositories
Muse Spark 1.1*
NEW61.50±3.10
gpt-5.4 (xHigh)*
59.10±3.56
Muse Spark*
55.00±3.60
Evaluating long-horizon software engineering tasks in commercial-grade private repositories
Muse Spark 1.1*
NEW51.50±5.50
claude-opus-4-6 (thinking)*
47.10±6.07
Muse Spark*
44.70±6.05
Forecasting scientific experiment outcomes
gemini-3-pro-preview
25.27±1.92
claude-opus-4-5-20251101
23.05±0.51
claude-opus-4-1-20250805
22.22±1.48
Challenging LLMs at the frontier of human knowledge
46.44±1.96
44.32±1.95
40.56±1.92
Challenging LLMs at the frontier of human knowledge
47.31±2.11
45.32±2.10
40.92±2.07
Evaluating spoken dialogue systems in multi-turn interaction
Inkling (Thinking)*
NEW56.64±4.55
Inkling-small
NEW54.87±4.57
gemini-3-pro-preview (Thinking)*
54.65±4.57
Evaluating spoken dialogue systems in multi-turn interaction
gpt-realtime-2 (xHigh)
48.45±4.59
tml-interaction-small
43.36±4.55
gpt-realtime-2
37.61±4.45
Evaluating spoken dialogue systems in multi-turn interaction
Inkling (Thinking)
NEW56.64±2.70
Inkling-small
NEW54.87±4.57
gemini-3-pro-preview (Thinking)
54.65±4.57
Evaluating Professional Reasoning in Finance
Muse Spark 1.1
55.01±0.14
claude-fable-5 (max)
NEW53.86±0.14
claude-opus-4-6 (Non-Thinking)
53.28±0.18
Evaluating Professional Reasoning in Legal Practice
Muse Spark 1.1
57.05±0.22
claude-fable-5 (max)
NEW52.56±0.54
Muse Spark
52.29±0.06
Evaluating AI agents ability to perform real-world, economically valuable remote work
Fable-5
15.80
Opus 4.8
8.33
Codex GPT 5.5
6.25
Simulating real-world pressure to choose between safe or harmful behavior
o3-2025-04-16
10.50±0.60
claude-sonnet-4-20250514
12.20±0.20
o4-mini-2025-04-16
15.80±0.40
Evaluating how LLMs can dynamically interact with and reason about visual information
Muse Spark 1.1
NEW44.77±2.82
gpt-5.4-2026-03-05 (reasoning effort = high)
29.17±0.13
gemini-3.1-pro-preview
28.97±0.91
Multilingual Native Reasoning Evaluation Benchmark for LLMs
Muse Spark 1.1
NEW65.59±2.87
65.20±1.24
64.74±2.88
Assessing models across diverse, interdisciplinary challenges
Muse Spark
75.52±4.05
Muse Spark 1.1
NEW75.30±0.60
gemini-3.1-pro-preview
71.37±1.74
Frontier Risk Evaluation for National Security and Public Safety
8.24±1.93
9.63±2.11
12.40±1.48
Evaluate model honesty when pressured to lie
96.28±0.41
96.13±0.57
Claude Sonnet 4 (Thinking)
95.33±2.29
Evaluating model performance on complex, multi-step reasoning tasks
claude-fable-5-high
39.28±2.80
gpt-5.6-sol-high
37.12±2.80
gemini-3.1-pro-preview-high
36.78±2.71
Vision-Language Understanding benchmark for multimodal models
Gemini 2.5 Pro Experimental (March 2025)
54.65±1.46
gemini-2.5-pro-preview-06-05
54.63±0.55
gpt-5.4-pro-2026-03-05
53.89±2.02
Evaluating model performance on common tutoring tasks for high school and AP-level subjects
Muse Spark
68.55±0.95
gpt-5.4-pro-2026-03-05
56.62±1.02
gemini-2.5-pro-preview-06-05
55.65±1.11
We conduct high-complexity evaluations to expose model failures, prevent benchmark saturation, and push model capabilities--while continuously evaluating the latest frontier models.
Humans design complex evaluations and define precise criteria to assess models, while LLMs scale evaluations--ensuring efficiency and alignment with human judgment.
Our leaderboards are built on carefully curated evaluation sets, combining private datasets to prevent overfitting and open-source datasets for broad benchmarking and comparability.
If you'd like to add your model to this leaderboard or a future version, please contact [email protected]. To ensure leaderboard integrity, we require that models can only be featured the FIRST TIME when an organization encounters the prompts.