close
Scale Labs

Research to Advance AI

Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

[SHOWDOWN]

Model-preference rankings from real-world usage.

1claude-opus-4-61071.55
1gpt-5.2-chat-latest1069.54
1claude-opus-4-7 (Thinking)1064.47
1claude-opus-4-71062.76
3gpt-5.5-2026-04-231053.02
View more

[PAPERS]

Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

Date Title
8/20/2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and Alignment
8/6/2026
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and Alignment
6/30/2026
DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?Agents, Enterprise, Evaluation and Alignment
6/29/2026
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsAgents, Evaluation and Alignment
6/19/2026
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld TasksAgents, Evaluation and Alignment
6/10/2026
Rubric-Guided Self-Distillation: Post-Training Without Rubric VerifiersPost-Training
View more

[BLOG]

Insights, analysis, and updates from Scale Labs