close
Scale Labs

[LEADERBOARDS]

testing the limits of AI.

Benchmarks for frontier, agentic, and safety capabilities

Benchmarks20+

Including benchmarks on agentic coding, frontier reasoning, and safety alignment.

Models evaluated100+

From leading AI labs including OpenAI, Anthropic, Google, Meta, and open-source contributors.

DrugDiscoveryBench

DrugDiscoveryBench: 82 expert-curated tasks evaluating how reliably frontier coding agents perform the computational work of early-stage drug discovery.

1

GPT-5.5 (mini-SWE-agent) xhigh

51.60±4.30

1

Claude Sonnet 5 (mini-swe-agent) max

50.04±0.70

1

Gemini 3.5 Flash (Gemini CLI) high

50.00±2.40

View Full Ranking

SWE Atlas - Refactoring

Evaluating an agent's ability to restructure code while preserving behavior.

1

Fable-5 (Claude Code) xHigh

54.76±6.76

1

Opus-4.7 (Claude Code)

48.57±6.73

1

Opus 4.8 (Claude Code)

46.67±6.75

View Full Ranking

SWE Atlas - Test Writing

Evaluating an agent’s ability to write production-grade tests

1

Opus 5 (Claude Code) xHigh

NEW

62.22±5.58

1

Fable-5 (Claude Code) xHigh*

NEW

55.60±5.80

1

Opus 4.8 (Claude Code) xhigh

49.63±5.93

View Full Ranking

SWE Atlas - Codebase QnA

Evaluating deep code comprehension and reasoning

1

Opus 5 (Claude Code) xHigh

NEW

63.17±5.01

1

Opus 4.8 (Claude Code) xhigh

57.26±4.93

2

GLM 5.2 (Mini-SWE-Agent)

48.12±5.07

View Full Ranking

HiL-Bench (Human-in-Loop Benchmark)

Evaluates whether agents recognize information gaps and ask targeted clarifying questions.

1

Claude Opus 5

NEW

57.00±5.48

1

Claude Fable 5

56.33±5.50

2

GLM 5.2

43.67±5.65

View Full Ranking

MCP Atlas

Evaluating real-world tool use through the Model Context Protocol (MCP)

1

Muse Spark 1.1

NEW

88.10±1.95

2

claude-opus-5 (xhigh)

NEW

85.80±2.10

2

gemini-3.5-flash (high)

83.60±2.30

View Full Ranking

SWE-Bench Pro (Public Dataset)

Evaluating long-horizon software engineering tasks in public open source repositories

1

Muse Spark 1.1*

NEW

61.50±3.10

1

gpt-5.4 (xHigh)*

59.10±3.56

3

Muse Spark*

55.00±3.60

View Full Ranking

SWE-Bench Pro (Private Dataset)

Evaluating long-horizon software engineering tasks in commercial-grade private repositories

1

Muse Spark 1.1*

NEW

51.50±5.50

1

claude-opus-4-6 (thinking)*

47.10±6.07

3

Muse Spark*

44.70±6.05

View Full Ranking

SciPredict

Forecasting scientific experiment outcomes

1

gemini-3-pro-preview

25.27±1.92

1

claude-opus-4-5-20251101

23.05±0.51

1

claude-opus-4-1-20250805

22.22±1.48

View Full Ranking

Humanity's Last Exam

Challenging LLMs at the frontier of human knowledge

1

46.44±1.96

1

44.32±1.95

3

40.56±1.92

View Full Ranking

Humanity's Last Exam (Text Only)

Challenging LLMs at the frontier of human knowledge

1

47.31±2.11

1

45.32±2.10

3

40.92±2.07

View Full Ranking

AudioMultiChallenge

Evaluating spoken dialogue systems in multi-turn interaction

1

Inkling (Thinking)*

NEW

56.64±4.55

1

Inkling-small

NEW

54.87±4.57

1

gemini-3-pro-preview (Thinking)*

54.65±4.57

View Full Ranking

AudioMultiChallenge - Audio Output

Evaluating spoken dialogue systems in multi-turn interaction

1

gpt-realtime-2 (xHigh)

48.45±4.59

1

tml-interaction-small

43.36±4.55

3

gpt-realtime-2

37.61±4.45

View Full Ranking

AudioMultiChallenge - Text Output

Evaluating spoken dialogue systems in multi-turn interaction

1

Inkling (Thinking)

NEW

56.64±2.70

1

Inkling-small

NEW

54.87±4.57

1

gemini-3-pro-preview (Thinking)

54.65±4.57

View Full Ranking

Professional Reasoning Benchmark - Finance

Evaluating Professional Reasoning in Finance

1

Muse Spark 1.1

55.01±0.14

2

claude-fable-5 (max)

NEW

53.86±0.14

2

claude-opus-4-6 (Non-Thinking)

53.28±0.18

View Full Ranking

Professional Reasoning Benchmark - Legal

Evaluating Professional Reasoning in Legal Practice

1

Muse Spark 1.1

57.05±0.22

2

claude-fable-5 (max)

NEW

52.56±0.54

2

Muse Spark

52.29±0.06

View Full Ranking

Remote Labor Index (RLI)

Evaluating AI agents ability to perform real-world, economically valuable remote work

1

Fable-5

15.80

2

Opus 4.8

8.33

3

Codex GPT 5.5

6.25

View Full Ranking

PropensityBench

Simulating real-world pressure to choose between safe or harmful behavior

1

o3-2025-04-16

10.50±0.60

2

claude-sonnet-4-20250514

12.20±0.20

3

o4-mini-2025-04-16

15.80±0.40

View Full Ranking

VisualToolBench (VTB)

Evaluating how LLMs can dynamically interact with and reason about visual information

1

Muse Spark 1.1

NEW

44.77±2.82

2

gpt-5.4-2026-03-05 (reasoning effort = high)

29.17±0.13

2

gemini-3.1-pro-preview

28.97±0.91

View Full Ranking

MultiNRC

Multilingual Native Reasoning Evaluation Benchmark for LLMs

1

Muse Spark 1.1

NEW

65.59±2.87

1

65.20±1.24

1

64.74±2.88

View Full Ranking

MultiChallenge

Assessing models across diverse, interdisciplinary challenges

1

Muse Spark

75.52±4.05

1

Muse Spark 1.1

NEW

75.30±0.60

3

gemini-3.1-pro-preview

71.37±1.74

View Full Ranking

Fortress

Frontier Risk Evaluation for National Security and Public Safety

1

8.24±1.93

1

9.63±2.11

3

12.40±1.48

View Full Ranking

MASK

Evaluate model honesty when pressured to lie

1

96.28±0.41

1

96.13±0.57

1

Claude Sonnet 4 (Thinking)

95.33±2.29

View Full Ranking

EnigmaEval

Evaluating model performance on complex, multi-step reasoning tasks

1

claude-fable-5-high

39.28±2.80

1

gpt-5.6-sol-high

37.12±2.80

1

gemini-3.1-pro-preview-high

36.78±2.71

View Full Ranking

VISTA

Vision-Language Understanding benchmark for multimodal models

1

Gemini 2.5 Pro Experimental (March 2025)

54.65±1.46

1

gemini-2.5-pro-preview-06-05

54.63±0.55

1

gpt-5.4-pro-2026-03-05

53.89±2.02

View Full Ranking

TutorBench

Evaluating model performance on common tutoring tasks for high school and AP-level subjects

1

Muse Spark

68.55±0.95

1

gpt-5.4-pro-2026-03-05

56.62±1.02

1

gemini-2.5-pro-preview-06-05

55.65±1.11

View Full Ranking

Frontier AI Model Evaluations & Benchmarks

We conduct high-complexity evaluations to expose model failures, prevent benchmark saturation, and push model capabilities--while continuously evaluating the latest frontier models.

Scaling with Human Expertise

Humans design complex evaluations and define precise criteria to assess models, while LLMs scale evaluations--ensuring efficiency and alignment with human judgment.

Robust Datasets for Reliable AI Benchmarks

Our leaderboards are built on carefully curated evaluation sets, combining private datasets to prevent overfitting and open-source datasets for broad benchmarking and comparability.

[EVALUATE YOUR MODEL]

If you'd like to add your model to this leaderboard or a future version, please contact [email protected]. To ensure leaderboard integrity, we require that models can only be featured the FIRST TIME when an organization encounters the prompts.