close

DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

Image 1
Comments
4 min read
Labs are ditching factual knowledge for reasoning speed

Labs are ditching factual knowledge for reasoning speed

Comments
2 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Comments
6 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Image 1
Comments
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering

Meta's Muse Code Clears 59% on Deep Software Engineering

Comments
2 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Comments
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

Comments
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

Comments
3 min read
AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Comments
2 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

Image Image 6
Comments 1
4 min read
Agent Leaderboards Measure Score. We Added Price.

Agent Leaderboards Measure Score. We Added Price.

Comments
5 min read
How I tried to write an article about slow Chinese LLMs

How I tried to write an article about slow Chinese LLMs

Image Image Image 17
Comments 16
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.