AWS released G7 instances on April 2024, saying they deliver better price-performance for small LLM inference workloads. It is the company's first hardware update in the generative AI inference category since the launch of G6 in 2023.

AWS reported 12.5% lower cost-per-token for G7 12xlarge configurations compared to G6 12xlarge, measured on a 128 input token, 128 output token workload. That compares with 15% higher cost-per-token for G6 12xlarge in the same scenario.

G7 instances are built on NVIDIA Blackwell GPUs and target AI coding assistants and enterprise AI assistants. Availability begins in US East (Ohio) and US West (Oregon) regions, initially for developers and enterprises.

"The new G7 instances offer a structural advantage for MoE deployment due to native FP4 Tensor Core support," said Alex Kuznetsov, Machine Learning Specialist at AWS. "This enables faster decoding and lower latency for complex workloads."

The announcement follows AWS's release of SageMaker AI Generative AI Inference Recommendations. AWS said the G7 instances demonstrate how new GPU generations can improve inference efficiency for small LLMs.

AWS did not say how G7 instances will perform with larger models, and raised the question of whether the benefits will extend to more complex workloads. The company said it will continue to benchmark and optimize GPU configurations for various use cases.

Source: awsml