
AWS and NVIDIA announced an expanded partnership on August 26, 2026, with a plan to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure in 2027 and 2028. This is a forward-looking deployment plan, not capacity that is already installed or immediately available. The announcement places agentic AI and physical AI inside the same infrastructure story.
The partnership is about more than the GPU count. NVIDIA says Vera CPUs, NVLink Fusion, and NVHBM will be brought into next-generation AWS systems so the CPU, accelerator, memory, and interconnect can be designed together. For an agent running over a long workflow, the bottleneck may be outside model inference: how often data moves, whether memory is sufficient, whether the network stays low-latency, and whether tools or sub-agents have to wait for one another. Full-stack design is aimed at those constraints.
On models and data software, the announcement says NVIDIA Nemotron open models will be available through Amazon Bedrock and SageMaker, while cuDF and cuVS will come to Amazon EMR and OpenSearch. The combination targets more than training. It is intended to put retrieval, data processing, inference, and workflow execution closer together on the cloud platform. For an agent, indexing and retrieval speed can affect task completion more directly than adding another layer of model parameters.
The partnership also covers physical AI. NVIDIA says AWS global infrastructure will support robots and other physical systems with its simulation, perception, training, and deployment tools. The companies also plan to provide 100,000 GPUs on secure AWS infrastructure for US federal and national-security workloads, targeting IL6+ environments. These are future plans and vendor statements; actual timing, supply, customer access, and compliance scope still need deployment evidence.
The 2 million figure is attention-grabbing, but the operational question is whether the capacity becomes usable work. Agent workflows still need data permissions, tool security, observability, versioning, cost controls, and human approvals. Physical AI adds the gap between simulation and the real world, equipment failures, safety boundaries, and on-site latency. More compute alone does not solve those operating problems.
At this scale, infrastructure also amplifies pressure on power, cooling, networks, and supply chains. The performance, demand, and market-benefit statements in the AWS and NVIDIA release are company claims or forward-looking statements, not realized independent measurements. The market will need to watch actual GPU rollout speed, model utilization, cost per successful task, and whether customers can keep using the stack under governance requirements.
The long-term signal is that agentic AI is moving from model-service competition toward infrastructure-combination competition. GPUs, CPUs, memory, interconnects, databases, open models, cloud control planes, and physical-AI tools have to operate together. Two million GPUs is a substantial supply commitment, but whether it supports more reliable and lower-latency agent workflows will depend on the whole stack in production.



