What's new:
- Braintrust's Eval library, a resource of open source evals and 20+ skills, for running your own evals and inspecting our results
- Kimi K3 and DeepSeek V4 join the Braintrust model lineup across playgrounds, prompts, and the Gateway, no AI provider setup required
-
There are two ways to score a coding agent. By its output, or by its behavior.
Output scoring tells you if the agent produced a desired result, but can miss cases where the agent does something it was instructed not to.
So we ran an eval to see if scoring by behavior can catch
Evals are hard. Good eval research should be accessible.
At Braintrust, we have unique access to how leading AI teams think about quality. We study the best existing research and conduct our own experiments.
Today we're making this work available in one place, with an open