OpenAI publishes first Jalapeño results as inference infrastructure targets interactive agents

OpenAI says its Jalapeño inference chip improved work per watt and end-to-end latency across several models and interactive workloads, although the results remain vendor-reported.

OpenAI published the first measured results for Jalapeño on August 25, 2026. It is a follow-up to the company's June announcement about developing custom AI chips with Broadcom, moving the story from what it planned to build to how the system performed in testing. OpenAI describes Jalapeño as an inference chip designed alongside memory, networking, and software for latency-sensitive interactive work, including agents that repeatedly call tools.

The company used the public InferenceX benchmark from SemiAnalysis to compare GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI reports 1.5 to 1.9 times more AI work per watt at peak load and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For interactive workloads, it reports 2.1 to 4.1 times higher performance. These numbers should be read as OpenAI's own measurements, not as an independently replicated industry benchmark.

The important point is that the test is not only about the chip. OpenAI says InferenceX measures prefill, decode, memory access, network transfer, and data movement as part of the user experience. An agent may need to retrieve data, call a tool, wait for a result, and generate its next instruction. Small improvements at each turn can compound across the workflow. The more useful operational metric is therefore the latency, energy, and cost required to complete a successful task, not isolated token speed.

OpenAI also says Jalapeño is rated at 700W but measured at no more than 550W sustained, while the comparison systems are listed at 1,200W and 1,400W. Power and performance comparisons depend on the configuration, model, software version, and workload. They do not imply that every model or production environment will see the same ratio. Real infrastructure gains will also depend on deployment density, cooling, network bottlenecks, utilization, and the number of tool calls in each job.

The design process is another signal. OpenAI says Jalapeño moved from architecture to tape-out in nine months with AI assisting parts of the design. It also reports that selected GPT-OSS hardware blocks were 1.5 to 1.8 times faster when AI-generated. That is a result for selected blocks, not for the whole chip or model. It points to a broader possibility: model companies may apply AI simultaneously to models, compilers, hardware design, and runtime optimization.

OpenAI plans to begin deploying Jalapeño in its compute infrastructure by the end of 2026. The system is still being qualified, its software is maturing, and it is being validated across more models. OpenAI says it will continue using NVIDIA and other accelerators, so this is not an immediate replacement story. The meaningful test will be whether the results hold across models, toolchains, and sustained operation rather than only at a launch benchmark.

Jalapeño is worth watching because it brings agent infrastructure metrics closer to real workflows: completion quality, turn-to-turn latency, energy per successful task, and dependable capacity during peaks. The main performance and deployment claims remain vendor-reported, however, with no public independent replication yet. The cautious conclusion is that OpenAI is trying to co-optimize inference hardware and interactive agent needs, while the production value still has to be demonstrated.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.