About the session
Everyone wants to compare AI agents, but few agree on what good performance actually means. Ahmed Bashir, Jeff Smith, and Brandon Grabowski discuss how DevRev is building fair, transparent benchmarks for AI agents and why responsible measurement begins with the evaluation process, not simply a score.
This session examines the differences between agent and model benchmarking, the limitations of marketing-driven leaderboards, and the importance of measuring reliability, transparency, and practical impact.
Key takeaways
- Understand why headline benchmark scores may not reflect real-world work.
- Recognize what makes AI agent benchmarking different from traditional model evaluation.
- Consider reliability, transparency, and impact alongside accuracy.
Agenda
Defining good agent performance
Why evaluation begins with a clear definition of success.
Agent benchmarks versus model benchmarks
How the evaluation problem changes when assessing agents.
Fair and transparent measurement
The role of evaluation process and methodology.
Metrics beyond accuracy
Reliability, transparency, and practical impact.
Speakers

Ahmed Bashir
Chief Technology Officer, DevRev
- JS
Jeff Smith
Wrangling agentic systems, benchmarking and leaderboards @ DevRev
- BG
Brandon Grabowski




