NVIDIA argues that agent evaluation must measure completed tasks, not just tool calls

NVIDIA’s agent-evaluation guidance separates step-level process scoring from end-to-end outcome scoring and recommends tracking success, consistency, tool precision, steps, and cost.

NVIDIA published a technical article on AI agent evaluation on September 21, 2026. Its central point is simple: an agent that sounds as if it finished a job may not have finished it. Once an agent operates in a live environment, makes sequential tool calls, handles errors, and changes state, a single answer or function-call score is not enough.

The article argues that an evaluation environment should execute each tool call, track state across steps, and inspect the final world state. A refund agent, for example, should not receive full credit merely for selecting the right refund API if it skipped required checks, failed to update records, or never caused the refund to post. Outcome-oriented evaluation is closer to what users experience than isolated call selection.

NVIDIA separates evaluation into two layers. Step-level process scoring asks whether a call was valid, relevant, and useful given the state at that moment. End-to-end outcome scoring checks only the final state, such as whether a ticket routed correctly or a database actually changed. The first layer helps debug where a chain broke; the second is the completion gate that matters in production.

Both layers should attach to the same trace. A trace is the ordered record of an attempt: the user input, every action, and the environment state when the task stops. Without it, a team knows only that a task failed, not whether the cause was a wrong tool, bad arguments, state drift, a permission denial, or lost context on a later step.

The article recommends metrics beyond a single accuracy number: task success rate asks whether the environment reached the goal; consistency measures the range across three to five trials; tool-call precision exposes hallucinated or redundant calls; argument accuracy separates the right API with the wrong payload; steps per success and cost per success capture the resources required for a successful outcome.

NVIDIA also warns that benchmarks that all claim to test tool calling may not be comparable. Task complexity, statefulness, error recovery, and verification methodology change what a score means. Executable checks—tests passing or a database state changing—are generally stronger than a reference answer or an LLM judge; judge scores remain provisional until calibrated against human ratings.

Contamination is another issue. An agent with web access can find answers during an evaluation, and a public dataset can be scraped into future training data. The article therefore points to private domain evaluations built around real APIs, policies, and business tasks in environments that cannot be searched. That does not eliminate bias, but it reduces the blind spot created by evaluating text without the real system.

For teams deploying AI workflows, the method turns “is the agent good?” into an engineering question. Define the goal state, record the trace, separate process from outcome, and only then compare models, harnesses, or prompts. An agent that looks impressive in a demo but cannot reliably complete the real task has not passed a production evaluation.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.