AI Infrastructure
AWS and NVIDIA Add 2 Million GPUs to Power the Era of Agentic AI and Physical AI
AWS and NVIDIA have announced a major expansion of their partnership to meet rapidly growing demand for AI infrastructure, with plans to add 2 million additional NVIDIA GPUs to AWS infrastructure between 2027 and 2028 — alongside deeper collaboration spanning GPUs, CPUs, networking, open models, data processing, and robotics.
AWS and NVIDIA Accelerate AI Infrastructure Expansion
Amazon Web Services (AWS) and NVIDIA announced a major strategic expansion of their partnership on August 26, 2026, aimed at meeting continuously growing global demand for AI infrastructure.
At the center of the announcement is a plan to add 2 million more NVIDIA GPUs to AWS's global infrastructure between 2027 and 2028, including next-generation NVIDIA GPUs such as Blackwell Ultra, Rubin, and Rubin Ultra, along with AI factories built to support large-scale AI workloads.
This investment comes as many organizations shift from AI pilots to running AI in production — across enterprise automation, agentic AI, scientific discovery, and robotics.
For businesses, this transition isn't just about choosing a capable AI model. It also requires attention to compute, networking, storage, data processing, security, and the ability to scale infrastructure all at once.
Moving From AI Pilots to Production Requires More Capable Infrastructure
In the early stages of AI adoption, many organizations start with small projects — chatbots, generative AI, document processing, or data analysis assistants. But as AI gets applied to more business processes, data volumes and workloads grow along with it.
For example, an AI system once used by a handful of teams might expand to organization-wide use, or an AI agent that originally just answered questions might evolve into a system that analyzes data, plans, and carries out multi-step tasks on a user's behalf.
This shift makes infrastructure a critical driver of AI success. AWS and NVIDIA see AI compute demand continuing to grow steadily, requiring infrastructure that can support large workloads both now and in the future.
2 Million More NVIDIA GPUs on AWS
The GPUs being deployed include NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra, designed to support the increasingly complex, high-performance AI workloads of the next generation.
AWS also plans to expand its NVIDIA Blackwell capabilities in several ways, including the NVIDIA RTX PRO 4500 Blackwell Server Edition for Amazon EC2 G7 instances.
For organizations, the ability to continuously add compute capacity matters, since AI workloads no longer stop at model training — they span data processing, training, inference, and real-time operation.
Infrastructure that supports large GPU fleets lets organizations scale AI workloads without having to redesign their systems entirely every time usage grows.
NVIDIA Vera CPU Is Coming to AWS
Another significant move is that AWS and NVIDIA are working together to bring the NVIDIA Vera CPU to AWS. Vera is designed for next-generation AI workloads, particularly those that need strong CPU performance alongside GPU acceleration.
This matters for agentic AI, since AI agents don't run on GPUs alone — they require multi-step processing such as data handling, tool calls, planning, and coordination across systems.
As a result, next-generation AI infrastructure is trending toward heterogeneous computing, combining CPUs, GPUs, custom silicon, and multiple types of accelerators. AWS has its own custom silicon in Trainium and Graviton, while NVIDIA has its own GPUs and CPUs. This interoperability between hardware types gives customers more flexibility to match infrastructure to each specific workload.
NVLink Fusion and the Connection to AWS Trainium
AWS and NVIDIA are also expanding their collaboration around NVIDIA NVLink Fusion, a high-speed chip interconnect technology for connecting compute components that need to work together in an AI system.
AWS previously announced support for NVLink Fusion in next-generation Trainium chips, and the two companies are now extending that collaboration to NVIDIA Custom High-Bandwidth Memory (NVHBM).
Increasing memory bandwidth matters for AI workloads because large models need to process and move enormous amounts of data between compute units and memory.
By combining NVLink Fusion with next-generation memory technology, AWS can build approaches that let Trainium and NVIDIA GPUs work together more efficiently within the same rack-scale architecture. For organizations, this reflects a broader shift in AI infrastructure — away from thinking of servers as standalone machines, and toward designing AI systems at the rack and cluster level.
AWS and NVIDIA Double Down on Agentic AI
One of the most closely watched technologies right now is agentic AI. Unlike traditional generative AI, which focuses on generating text or content from a prompt, agentic AI is designed to carry out multi-step work by planning, calling tools, and connecting with other systems.
Examples of enterprise use cases include:
- AI agents for customer service
- Systems that help manage internal documents and data
- AI for business data analysis
- Automation for repetitive workflows
- AI that supports IT operations
- AI that connects with business applications and enterprise systems
As agentic AI becomes more sophisticated, the infrastructure behind it has to support high volumes of both compute and data processing. Expanding GPU, CPU, and networking capacity is a key part of how AWS and NVIDIA support AI agents that need to run continuously and serve large numbers of users.
NVIDIA Nemotron Adds an Open Model Option
Another part of the partnership is support for NVIDIA Nemotron open models on AWS, available through Amazon Bedrock as managed and serverless models, as well as through Amazon SageMaker for customers who want to deploy and fine-tune models on their own infrastructure.
Having open models as an option matters because businesses have different requirements around data privacy, model customization, cost, and deployment.
Some organizations may prefer a managed AI service for speed, while others with specific data requirements or dedicated AI engineering teams may want more control over the model and infrastructure. Supporting both approaches lets organizations choose a deployment method that fits their own business requirements.
Accelerating Data Processing and Vector Indexing With GPUs
AI can't run efficiently without good data infrastructure. AWS and NVIDIA are extending their collaboration into data processing, bringing NVIDIA CUDA-X libraries such as cuDF and cuVS to AWS services.
Another key area is vector indexing on Amazon OpenSearch Service, which matters for AI systems built on semantic search and retrieval-augmented generation (RAG). As vector databases scale to billions of records, index creation and tuning can become a system bottleneck.
This is especially important for organizations building enterprise AI, since AI performance doesn't depend on the model alone — it also depends on how quickly data can be accessed and processed.
Physical AI and Robotics Are Moving Into Real-World Use
Beyond AI software, AWS and NVIDIA are also expanding their partnership into Physical AI — AI that can perceive and operate in the physical world. Amazon Robotics is working with NVIDIA to develop next-generation robots using technologies such as NVIDIA Jetson, NVIDIA Omniverse, and NVIDIA Isaac.
Modern robotics development requires more than hardware — it also depends on simulation, synthetic data, robot training, route optimization, and real-to-sim validation. GPU-accelerated Amazon EC2 is being used as part of the infrastructure supporting these workloads.
This approach can be applied across industries such as warehouse automation, manufacturing, logistics, and future robotics systems. Physical AI is emerging as another growth market alongside generative AI and agentic AI.
AI Infrastructure Needs to Prioritize Security and Reliability
For large organizations, adding compute capacity alone isn't enough — real-world AI workloads also require security, reliability, and strong network performance.
AWS and NVIDIA are focusing on integrating the NVIDIA platform with the AWS Nitro System and Elastic Fabric Adapter (EFA). The Nitro System supports security and isolation across AWS infrastructure, while EFA helps systems scale workloads efficiently across compute nodes.
These capabilities matter greatly for AI training and inference, which rely on large numbers of GPUs working together — particularly for organizations applying AI to sensitive data such as customer information, financial data, business data, or internal systems, where infrastructure security has to be considered from the design stage.
AI Infrastructure for Government Agencies
AWS and NVIDIA also announced plans to build AI factories for the U.S. government, with plans to provide 100,000 NVIDIA GPUs on AWS Secure Infrastructure to support federal and national security workloads.
This infrastructure will support workloads with high-level security requirements, including use cases involving Impact Level 6 or higher.
This underscores that AI infrastructure is becoming strategic infrastructure — not just a technology for experimentation or productivity gains. The private sector, government agencies, and large enterprises alike need infrastructure capable of running AI in production with appropriate security and compliance.
What Organizations Should Prepare For as AI Infrastructure Expands
The AWS-NVIDIA announcement makes clear that AI is entering a phase that demands continuously larger infrastructure. For organizations planning AI investments, the key consideration isn't just which AI model to choose — it's the full infrastructure picture, end to end.
1. Compute Capacity
Assess what type of CPU, GPU, or accelerator your workload needs, and how much future scaling it will need to support.
2. Data Infrastructure
Check the readiness of your data pipeline, storage, databases, and data management systems, since good AI depends on data that's ready to use.
3. Networking
Large AI clusters need high-bandwidth, low-latency networks so GPUs and compute nodes can work together efficiently.
4. Security
Data used for AI may be sensitive to the organization, so security, access control, and data protection need to be designed appropriately.
5. Scalability
Infrastructure should be able to scale as AI workloads grow, without requiring a complete architectural overhaul.
6. Cost Efficiency
Organizations should look at total cost of ownership (TCO), not just hardware or cloud resource pricing, since energy, networking, storage, and operations costs all affect overall AI cost.
Conclusion: AI Growth Depends on Ready Infrastructure
This partnership expansion between AWS and NVIDIA shows that the industry is entering an era where AI requires larger, more complex infrastructure than before. Adding 2 million more NVIDIA GPUs between 2027 and 2028 is just one part of a broader ecosystem expansion spanning GPUs and CPUs to networking, open models, data processing, and robotics.
At the same time, the rise of agentic AI and physical AI is reshaping what AI workloads look like — shifting from a focus on content generation toward systems that can analyze, plan, decide, and act.
For organizations planning digital transformation or AI transformation, this is an important signal: preparing AI infrastructure shouldn't be an afterthought — it should be part of IT strategy from the start.
Because as AI shifts from pilot to production, an organization's ability to scale AI may come down not to whether it "has AI," but to how ready its infrastructure is to grow alongside the business.