NVIDIA’s SWE-Serve shows why coding agents need live-serving tests

NVIDIA’s SWE-Serve benchmark evaluates coding-agent patches on 53 inference-serving tasks and shows how local checks can overestimate real serving reliability.

NVIDIA introduced SWE-Serve on September 23, 2026, as an inference-serving benchmark for AI coding agents. It does not only ask whether an agent can edit a repository and pass local checks. It tests whether the patch still works when a real model server starts, loads the model, and handles requests.

SWE-Serve contains 53 tasks, including 19 tasks that load the required model and test the agent’s patch through a live serving interface. The complete verifier contains 276 live-serving tests covering model loading, expert routing, text and image serving, batched generation, output ordering, and log probabilities.

NVIDIA reports that 627 patches pass 45.9% of the time under the complete verifier. When the live-serving tests are removed, the rate rises to 69.4; 147 patches change from fail to pass when the real serving path is excluded. The gap shows how much local checks can miss.

Tasks that cross runtime domains are harder as well. The 26 tasks confined to one domain have a 69.0% pass rate under the best tested settings, while 27 tasks spanning request handling, scheduling, model execution, and KV-cache or runtime-resource management reach 47.7%. That is a 21.3-point difference.

Across 11 models and 31 model-effort configurations, the best pass@1 means range from 34.6% to 75.5%. No model leads every engineering family. NVIDIA also stresses that a SWE-Serve pass only means that the patch satisfied the benchmark verifier; it does not establish deployability, merge readiness, or endorsement by upstream maintainers.

To protect evaluation integrity, SWE-Serve runs closed-book evaluations that block the public web and upstream source repositories, while allowing access to Hugging Face model weights. NVIDIA audited 1,749 trials and says 196 prohibited retrieval attempts were blocked, with none succeeding. The design prevents an agent from simply retrieving task-specific upstream code.

For teams adopting coding agents, the lesson is to separate levels of evidence: static checks, local unit tests, server startup, real requests, multi-domain behavior, cost, and time each need a gate. An agent can produce a patch quickly, but deployment still requires end-to-end verification in an environment close to where the system will run.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.