NVIDIA · Q2 2023

HGX Rubin NVL8

The NVIDIA HGX Rubin NVL8 is a high-performance GPU module designed for datacenter environments, targeting AI training and high-performance computing workloads. It is part of NVIDIA's Hopper architecture, offering significant advancements in compute capabilities and memory bandwidth. The NVL8 variant is optimized for large-scale deployments, providing exceptional scalability and efficiency.

HGX Rubin NVL8 — illustration of the card's form factor
VRAM
288 GB
FP32 TFLOPS
130 TFLOPS
CUDA Cores
16,896
Memory
HBM4

Compute Performance

FP6433 TFLOPS
FP32130 TFLOPS
TF322 PFLOPS TFLOPS
FP164 PFLOPS TFLOPS
BF164 PFLOPS TFLOPS
FP817.5 PFLOPS TFLOPS

Architecture

MicroarchitectureRubin
Sparse AccelerationSupported (Sparse NVFP4)
Dynamic PrecisionSupported (NVFP4, FP8/FP6, FP16/BF16, TF32, FP32, FP64, INT8)

Memory & VRAM

Memory TypeHBM4
Total Capacity288 GB
Bandwidth22 TB/s

Connectivity & Scaling

InterconnectNVIDIA NVLink
GenerationSixth Generation
IB Bandwidth3.6 TB/s
TopologyNVLink Switch
Max GPUs/Node8
Scale-OutNVIDIA Quantum InfiniBand; Spectrum-X Ethernet; Spectrum-XGS Ethernet; NVIDIA BlueField DPU; DOCA

Physical Design

Form FactorSXM

Server & Deployment

OEM AvailabilityCPU and memory specs are defined by OEM offerings.
Preconfigured["NVIDIA DGX Rubin NVL8 (turnkey infrastructure solution)","NVIDIA DGX SuperPOD","NVIDIA DGX BasePOD"]
DGX/HGXNVIDIA HGX Rubin NVL8 can be configured as HGX Vera Rubin NVL8 (when paired with Vera CPUs) or with x86-based CPU baseboards.
Rack-Scale["NVIDIA DGX SuperPOD","NVIDIA DGX BasePOD"]
Ref Architectures["NVIDIA DGX BasePOD","NVIDIA DGX SuperPOD"]

System Compatibility

CPU PairingCan be paired with an NVIDIA Vera CPU or an x86-based baseboard.
MotherboardAvailable in a single baseboard with eight NVIDIA Rubin SXM GPUs (supports Rubin, Blackwell, or Blackwell Ultra SXMs).

Benchmarks & Throughput

Structured Sparsity

NVFP4 Inference specification is sparse.

Transformer Throughput

With architectural innovations including 400 PFLOPS of NVFP4 compute, 3x more memory bandwidth at 176 TB/s, and 2x more NVLink Switch bandwidth at 28.8 TB/s for high-throughput inter-GPU communication, HGX Rubin NVL8 delivers 10x more token factory throughput versus HGX B200.

Training Benchmarks

NVFP4 Training2 | 35 PFLOPS; FP8/FP6 Training2 | 17.5 PFLOPS

Inference Benchmarks

NVFP4 Inference | 50 PFLOPS

Scaling Efficiency

HGX Rubin NVL8 brings breakthrough mixture-of-experts pretraining to the 8 GPU server form factor, training next-generation agentic AI models with 4x fewer GPUs, enabled by architectural innovations including 4x more NVFP4 training FLOPS, 1.6x more high-speed HBM memory capacity, and 2x more NVLink bandwidth versus HGX B200.

Multi-GPU Scalability

Scaling Characteristics

Network Bottlenecksdeterministic latency, lossless throughput, stable iteration times, and the ability to scale not only within a data center but also across multiple sites.
Parallelismleverages sixth-generation NVIDIA NVLink to ensure seamless peer-to-peer communication for massive model parallelism.

Workload Readiness

LLM Training

35 PFLOPS

LLM Inference

50 PFLOPS

Vision Training

4 PFLOPS

Reinforcement Learning

dedicated reinforcement learning engine that optimizes memory movement in hardware

HPC / Simulation

33 TFLOPS

Scientific Computing

200 TFLOPS

Real-Time Serving

50 PFLOPS

Market Authority

Cloud Adoption

Vera Rubin ramping into production and shipping systems to hyperscalers

Key Strengths

This GPU excels at large-scale AI training and inference tasks, offering superior performance in deep learning frameworks. Its architecture is optimized for high throughput and low latency, making it ideal for complex simulations and scientific computing. The NVL8's scalability and efficiency make it a standout choice for demanding datacenter applications.

Limitations

While the HGX Rubin NVL8 offers exceptional performance, its high power requirements and need for advanced cooling solutions can be a trade-off for some deployments. Additionally, its availability may be limited due to high demand and production constraints, potentially impacting procurement timelines for large-scale projects.

Expert Insight

The HGX Rubin NVL8 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.