About the session
Everyone's racing to build AI agents. Almost nobody agrees on how to measure them.
That's the problem Enterprise-Bench was designed to solve - and on the next DevRev Live, we're pulling it apart piece by piece.
We'll start with why existing benchmarks fall short, then walk through L1 to L4 (our answer to what "good" actually looks like at each level of complexity - and how it stacks up against TPC).
From there, we'll dig into the types of queries we chose, why the dataset looks the way it does, and how we're judging results without fooling ourselves.
We'll also sit down with Alex Dimakis from Bespoke Labs to zoom out - what's happening across the industry, why this approach is different, and what we're chasing after release.
If you care about AI agents that actually work in production, this one's for you.
Speakers

Ahmed Bashir
Chief Technology Officer, DevRev
- JS
Jeff Smith
Wrangling agentic systems, benchmarking and leaderboards @ DevRev

Alex Dimakis
UC Berkeley Prof, Co-founder & Chief Scientist




