The evaluation gap

Traditional benchmarks measure single-turn question answering, which says little about whether an agent can reliably complete a multi-step task involving tool use and recovery from errors.

New approaches

Emerging benchmarks measure task completion rate across long-horizon workflows, including how gracefully an agent handles a failed tool call or ambiguous instruction rather than just final-answer accuracy.

Why it matters

As agents take on more autonomous, multi-step work, evaluation methodology needs to catch up — a model that scores well on QA benchmarks can still fail badly in a real agentic deployment.