Amazon Web Services (AWS) and NVIDIA have announced a major expansion of their cloud infrastructure partnership, committing to deploy two million additional high-end NVIDIA GPUs across AWS global data centers in 2027 and 2028. The deployment expands on AWS's previous commitment from GTC 2026 to add one million GPUs starting in 2026, bringing total forward allocations across the multi-year cycle to three million units.
The upcoming capacity will comprise NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra architectures. The hardware is designated to support frontier training, large-scale agentic AI systems, enterprise automation, and physical AI workloads.
Hardware Architecture and Next-Generation Deployments
The expanded infrastructure program introduces several architectural shifts across compute nodes, memory topologies, and networking fabrics:
- Vera CPU Integration: AWS will introduce instances powered by NVIDIA's standalone ARM-based Vera CPUs, designed to work alongside next-generation Rubin accelerators.
- Rubin and Rubin Ultra Silicon: The Vera Rubin platform succeeds Blackwell, delivering up to 50 petaflops of FP8 inference per accelerator when paired with Vera processors and supporting up to 288 GB of high-bandwidth memory (HBM4). Rubin systems will deploy in rack-scale NVL144 configurations.
- NVLink Fusion with NVHBM: The deployment will incorporate NVIDIA NVLink Fusion interconnects featuring custom NVIDIA high-bandwidth memory (NVHBM) architectures.
- Nitro System and EFA Binding: Accelerators will interface with AWS Nitro System virtualization engines and Elastic Fabric Adapter (EFA) networking to provide line-rate throughput and isolation across multi-tenant partitions.
- Workstation Expansion: AWS is also expanding Blackwell capacity for Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs.

Dedicated AI Factories and Federal Capacity
As part of the expanded agreement, AWS and NVIDIA will construct specialized "AI factories" tailored for enterprise and public-sector compute demands.
The initiative includes provisioning a dedicated pool of 100,000 GPUs deployed across sovereign, secure AWS infrastructure specifically designated for United States government agencies and national security research workloads. These isolated clusters will support classified and restricted data processing pipelines while operating under federal compliance frameworks.
Cloud Allocation Visibility
Securing guaranteed delivery schedules for two million advanced accelerators across 2027 and 2028 provides AWS with multi-year hardware visibility in an environment where advanced packaging and high-bandwidth memory constraints continue to dictate hyperscaler capacity limits. By locking in Blackwell Ultra and Vera Rubin silicon allocations early, AWS aims to ensure continuous instance availability as client model parameters scale into tens of trillions.



