I spend a lot of time on benchmarks, these days. I help maintain several public benchmarks, work with customers on private benchmarks, and talk to partners about how we can all build benchmarks that are useful for evaluating all of the components of our voice agents.
This
Introducing VAmoS Bench from @veris_ai 🔥
3,300 calls, 11 agents, 100 scenarios
We compared what agents said with what they did
@pipecat_ai @trydaily @livekit @Vapi_AI @elevenlabs @cartesia @OpenAI @GoogleAI @retellai @NVIDIAAI @DeepgramAI
veris.ai/leaderboard
🧵








