Amazon Web Services (AWS) and NVIDIA have significantly expanded their strategic partnership, committing to deploy 2 million additional NVIDIA Graphics Processing Units (GPUs) across AWS’s worldwide infrastructure between 2027 and 2028. This major joint effort aims to deepen computational capabilities and deliver comprehensive, co-engineered solutions that will help clients accelerate advanced Artificial Intelligence (AI) development at an unprecedented scale.
Strategic Expansion and Collaboration Scope
Building upon a nearly two-decade history of joint innovation, the companies plan to enhance their collaboration far beyond GPU deployment. Their partnership will cover key areas including CPUs, networking, open models, data processing, and robotics, establishing comprehensive “AI factories” for various sectors.
The increased capacity is designed to complement AWS’s proprietary custom silicon, such as Trainium chips. This provides customers with the flexibility to choose the optimal compute resources for their unique operational requirements, allowing them to utilize NVIDIA GPUs, AWS Trainium chips, or a combination of both.
AWS CEO Matt Garman emphasized the importance of this flexible infrastructure, stating, “Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together. That’s why we’ve invested deeply with NVIDIA to make AWS the best place to run NVIDIA AI technologies, optimizing performance across our infrastructure from networking and security to deployment. This expanded collaboration gives frontier labs, enterprises, and governments even more ways to build and deploy AI on AWS.”
Similarly, NVIDIA founder and CEO Jensen Huang highlighted the scale of the commitment, noting, “For 16 years, we have scaled NVIDIA computing in the cloud together. Now we are expanding our partnership across the full stack—GPUs, CPUs, networking, open models and software—to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver. This expansion reflects customers’ demand for NVIDIA’s platform on AWS.”
Advanced Compute and Connectivity
The expansion is driven by the rapid scaling of AI workloads, which range from model training and deployment to data processing used in intelligent applications. As customers advance from initial pilot projects to full-scale production in areas like agentic AI, scientific discovery, enterprise automation, and robotics, they require robust, scalable infrastructure.
In addition to the GPU commitment, the partnership is introducing new compute options for specialized AI tasks. AWS and NVIDIA are collaborating to bring Vera CPU-based infrastructure to AWS. This offering provides an alternative for agentic AI workloads that require high-performance CPU computing alongside accelerated resources. Vera is particularly suited for the CPU-intensive tasks foundational to agentic AI and reinforcement learning, such as code execution, tool utilization, sandboxing, and data pipeline orchestration.
Furthermore, AWS is enhancing its silicon integration. At the re:Invent 2025 event, AWS announced support for NVIDIA’s NVLink Fusion high-speed chip interconnect technology in next-generation Trainium chips. This integration allows Amazon’s Annapurna Labs to leverage NVIDIA’s custom high-bandwidth memory (NVHBM) technology, leading to enhanced performance and efficiency for AI workloads while seamlessly connecting Trainium and GPUs within a standard rack architecture.
Government and Security Infrastructure
Addressing national security requirements, AWS and NVIDIA are planning to construct secure AI factories specifically for the U.S. government. This initiative will deliver NVIDIA’s complete AI stack, including 100,000 GPUs, running on AWS’s highly secure infrastructure. This capacity is designated for federal and national-security workloads classified at Impact Level 6 (IL6) and higher, representing one of the highest government security classifications.
Current Capabilities and Services
Existing deep technical integrations are already providing tangible benefits to customers today, including:
- Security and Reliability: All EC2 instances powered by NVIDIA GPUs or Trainium chips—including those utilizing NVLink Fusion—are built on the AWS Nitro System and connected via the Elastic Fabric Adapter (EFA), ensuring high security, reliability, and network performance.
- Open Models: The Nemotron family of open models from NVIDIA is accessible via Amazon Bedrock as fully managed, serverless models, and also through Amazon SageMaker for customers needing self-managed deployment and fine-tuning.
- Data Processing: GPU-accelerated data processing on Amazon EMR using NVIDIA cuDF can achieve up to 3.7 times faster processing speeds and 30% better price-performance for Apache Spark workloads compared to CPU-based setups.
- Vector Indexing: Amazon OpenSearch Service offers GPU-accelerated vector indexing, delivering up to 9 times faster indexing at a cost that is one-quarter of the traditional expense.
- Robotics: Amazon Robotics is integrating NVIDIA’s full-stack physical AI platform, which includes Jetson, Omniverse, and Isaac, to expedite next-generation warehouse automation through simulation, synthetic data generation, and real-world validation.