The hard truth about AI agent benchmarks

The hard truth about AI agent benchmarks

About the session

Everyone wants to compare AI agents, but few agree on what good performance actually means. Ahmed Bashir, Jeff Smith, and Brandon Grabowski discuss how DevRev is building fair, transparent benchmarks for AI agents and why responsible measurement begins with the evaluation process, not simply a score.

This session examines the differences between agent and model benchmarking, the limitations of marketing-driven leaderboards, and the importance of measuring reliability, transparency, and practical impact.

Key takeaways

  • Understand why headline benchmark scores may not reflect real-world work.
  • Recognize what makes AI agent benchmarking different from traditional model evaluation.
  • Consider reliability, transparency, and impact alongside accuracy.

Agenda

  1. Defining good agent performance

    Why evaluation begins with a clear definition of success.

  2. Agent benchmarks versus model benchmarks

    How the evaluation problem changes when assessing agents.

  3. Fair and transparent measurement

    The role of evaluation process and methodology.

  4. Metrics beyond accuracy

    Reliability, transparency, and practical impact.

Speakers

  • Ahmed Bashir

    Ahmed Bashir

    Chief Technology Officer, DevRev

  • JS

    Jeff Smith

    Wrangling agentic systems, benchmarking and leaderboards @ DevRev

  • BG

    Brandon Grabowski