NVIDIA and AWS Collaborate to Bring AI to Production at Scale

77d ago · US · primary source: blogs.nvidia.com

NVIDIA and Amazon Web Services are expanding their collaboration with new EC2 G7 instances and GPU-accelerated vector search in OpenSearch Serverless, aiming to streamline production-scale AI deployment for enterprise customers [1]. The new Amazon EC2 G7 instances are powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, targeting AI inference, graphics, video, and data analytics workloads [1]. Compared with the previous-generation G6 instances, the G7 delivers up to 4.6x AI inference performance and up to 2.1x graphics performance, according to NVIDIA [1]. The instances support configurations with up to eight GPUs, 256GB of total GPU memory, 700 Gbps of EFA-enabled networking, and up to 7.6TB of local NVMe SSD storage [1]. NVIDIA, headquartered in Santa Clara, California, originally focused on GPUs for video games before broadening into AI and high-performance computing [2]. The company now controls more than 80% of the market for GPUs used in training and deploying AI models [2]. On the retrieval side, the NVIDIA cuVS library is now the default compute choice for vector indexing in Amazon OpenSearch Serverless [1]. This integration makes GPU-powered vector search a standard capability rather than a specialized optimization project [1]. NVIDIA states that vector indexing is up to 10x faster and costs a quarter of the price compared with CPU-only builds [1]. The shift enables teams building retrieval-augmented generation, semantic search, and agentic AI applications to construct billion-scale vector databases in under an hour [1]. AWS has also achieved NVIDIA Exemplar Cloud status for the NVIDIA GB300 on training workloads, signaling that the cloud provider meets NVIDIA's performance benchmarks for large-scale AI training [1]. The designation stems from co-engineering efforts between the two companies and is intended to help developers evaluate cloud infrastructure with greater confidence [1]. AWS has invested heavily in custom silicon for its cloud, including the Arm-based Graviton processor family first announced in 2018, though the latest AI-focused instances rely on NVIDIA GPUs [5]. The broader AI infrastructure market has seen rapid growth, with NVIDIA becoming the first company to surpass $5 trillion in market capitalization in 2025, driven by demand for AI data center hardware [2].

infrastructure

Background sources we checked (5)
  • en.wikipedia.org ↗ Nvidia Corporation ( en-VID-ee-ə) is an American multinational technology company headquartered in Santa Clara, California. The company develops graphics processing units (GPUs), systems on chips (SoCs), and application programming interfaces (APIs) for data science, high-perform…
  • en.wikipedia.org ↗ OpenAI is an American artificial intelligence (AI) research organization headquartered in San Francisco, consisting of OpenAI Group PBC, a for-profit public benefit corporation (PBC), partially controlled by OpenAI Foundation, a nonprofit. OpenAI developed the generative pre-trai…
  • en.wikipedia.org ↗ Mirantis Inc. is a Campbell, California, based B2B open source cloud computing software and services company. Its primary container and cloud management products, part of the Mirantis Cloud Native Platform suite of products, are Mirantis Container Cloud and Mirantis Kubernetes En…
  • en.wikipedia.org ↗ AWS Graviton is a family of 64-bit Arm-based central processing units designed by Amazon Web Services (AWS) for use in its cloud computing infrastructure. The processors are part of AWS's custom silicon program and are used in Amazon Elastic Compute Cloud (EC2) instance families …
  • en.wikipedia.org ↗ Elemental was an American software company based in Portland, Oregon, and active from 2006 to 2015. It was founded by three engineers formerly of the semiconductor company Pixelworks: Sam Blackman (CEO), Jesse Rosenzweig (CTO), and Brian Lewis. In 2015, it was acquired by Amazon.…

Sources

Spot something wrong? Report an issue