shiichan

AI Inference Up to 4.6x Faster: Amazon EC2 G7 Instances Are Here!

Hey everyone, it's Shii-chan! Today I found some news that GPU fans are going to love. AWS has made its new GPU instance, Amazon EC2 G7, generally available!

AWS Blog aws.amazon.com

What was announced?

On the AWS Blog, AWS announced the general availability of Amazon EC2 G7 instances. G7 is a new-generation GPU instance powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, built to speed up workloads like AI inference, graphics, and data analytics. AWS also says it's the first major cloud provider to offer these GPUs as a managed service, so you're getting very fresh hardware here.

The story so far

Earlier G-series instances (the G6 generation) could already use GPUs, but the demands of AI inference, real-time graphics, and video processing kept climbing, and more workloads needed extra power. For generative AI inference and rendering in particular, GPU memory bandwidth and network speed translate directly into how fast things run, so a generational jump makes a real difference.

What changes

Compared to G6, here's how much faster G7 gets:

  • Up to 4.6x faster AI inference
  • Up to 2.1x better graphics performance
  • 1.33x more GPU memory capacity, with 2.45x the bandwidth
  • 7x higher network throughput (700 Gbps with EFA)
  • 1.5x more concurrent video encode/decode streams

Whether you run inference, drive 3D rendering or VDI (virtual desktops) in the cloud, or transcode video, there's something here for you.

Let's dive deeper

Let's start with the GPU. Each instance uses the NVIDIA RTX PRO 4500 Blackwell Server Edition, with 32 GB of GPU memory per unit. It has 5th Gen Tensor Cores, 4th Gen RT Cores, and ninth-generation NVENC and sixth-generation NVDEC engines, so it covers AI, ray tracing, and video processing all at once.

There are seven instance sizes to choose from:

Size GPUs GPU memory vCPUs Memory Local storage Network
g7.2xlarge 1 32 GB 8 32 GiB 600 GB NVMe 60 Gbps
g7.4xlarge 1 32 GB 16 64 GiB 600 GB NVMe 100 Gbps
g7.8xlarge 1 32 GB 32 128 GiB 950 GB NVMe 100 Gbps
g7.12xlarge 2 64 GB 48 192 GiB 1.9 TB NVMe 175 Gbps
g7.24xlarge 4 128 GB 96 384 GiB 3.8 TB NVMe 350 Gbps
g7.48xlarge 8 256 GB 192 768 GiB 7.6 TB NVMe 700 Gbps
g7.metal 8 256 GB 192 768 GiB 7.6 TB NVMe 700 Gbps

On the multi-GPU sizes, you can connect GPUs directly with NVIDIA GPUDirect P2P, and use NVIDIA GPUDirect RDMA together with EFA. It also works well with Amazon FSx for Lustre when you need fast reads and writes over large datasets. Supported operating systems include Amazon Linux, Ubuntu, RHEL, and Windows Server, and for graphics you get DirectX, Vulkan, and OpenGL support.

The use cases are broad: AI inference, graphics rendering, video transcoding and analytics, spatial computing, VDI, GPU-accelerated data analytics, and EMR or EKS workloads.

As for availability, G7 launches in US East (Ohio) and US West (Oregon) first. You can buy it On-Demand, with Savings Plans, or as Spot Instances, and the 12xlarge, 24xlarge, and 48xlarge sizes also support Dedicated Instances. (g7.metal is still on the way.)

Wrap-up

  • Amazon EC2 G7 is a new-generation GPU instance powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, and it's now generally available
  • Versus G6, it delivers up to 4.6x faster AI inference, up to 2.1x better graphics, and 7x higher network throughput (700 Gbps)
  • Each GPU has 32 GB of memory, and there are seven sizes from g7.2xlarge to g7.48xlarge
  • It's available in US East (Ohio) and US West (Oregon), via On-Demand, Spot, and more

If you want faster AI inference, or you run heavy graphics or video workloads in the cloud, this is a machine worth checking out!