The world's first benchmark for enterprise AI agents.

The world's first benchmark for enterprise AI agents.

About the session

Everyone's racing to build AI agents. Almost nobody agrees on how to measure them.
That's the problem Enterprise-Bench was designed to solve - and on the next DevRev Live, we're pulling it apart piece by piece.

We'll start with why existing benchmarks fall short, then walk through L1 to L4 (our answer to what "good" actually looks like at each level of complexity - and how it stacks up against TPC).

From there, we'll dig into the types of queries we chose, why the dataset looks the way it does, and how we're judging results without fooling ourselves.

We'll also sit down with Alex Dimakis from Bespoke Labs to zoom out - what's happening across the industry, why this approach is different, and what we're chasing after release.

If you care about AI agents that actually work in production, this one's for you.

Speakers

  • Ahmed Bashir

    Ahmed Bashir

    Chief Technology Officer, DevRev

  • JS

    Jeff Smith

    Wrangling agentic systems, benchmarking and leaderboards @ DevRev

  • Alex Dimakis

    Alex Dimakis

    UC Berkeley Prof, Co-founder & Chief Scientist