NVIDIAData Center GPU
NVIDIA H200 SXM
Based on Hopper architecture. 141GB memory with 700W TDP.
FP16 Performance
989
TFLOPS
Memory Capacity
141
GB
Power Draw
700
Watts TDP
Performance per Dollar
0.03 TFLOPS / $
Based on list price
Subject to market fluctuation
Benchmark Results
Published results grouped by model variant and scenario definition
18 resultsOpen Category View
Mistral 7B Instruct v0.3
FP8 Dense Serve · FP8
text-generation · vLLM · Batch 16 · Seq 4096 · 4096 in / 256 out
throughput
5,900 tokens/s
Latency P50
58 ms
Throughput
5,900 tokens/s
Power
670 W
Dense Mistral serving profile on NVIDIA H200.
Paraformer Large
INT8 Streaming ASR · INT8 · INT8
streaming-asr · FunASR · Batch 2 · 8s stream
latency
61 ms
Latency P50
61 ms
Throughput
18 ms
Power
600 W
Paraformer streaming profile on NVIDIA H200.
Wav2Vec2 Large 960h
FP16 Offline ASR · FP16
speech-recognition · ONNX Runtime · Batch 8 · 20s clip
throughput
42 x realtime
Latency P50
82 ms
Throughput
42 x realtime
Power
590 W
Wav2Vec2 offline ASR profile on NVIDIA H200.
ResNet-50
INT8 224x224 · INT8 · INT8
image-classification · OpenVINO · Batch 256 · 224x224
throughput
18,400 images/s
Latency P50
9 ms
Throughput
18,400 images/s
Power
600 W
ResNet-50 profile on NVIDIA H200.
RT-DETR-L
FP16 960x960 · FP16
object-detection · TensorRT · Batch 8 · 960x960
throughput
890 images/s
Latency P50
15 ms
Throughput
890 images/s
Power
650 W
RT-DETR profile on NVIDIA H200.
Gemma 2 9B Instruct
INT4 Chat Serve · BF16 · INT4
text-generation · TensorRT-LLM · Batch 4 · Seq 2048 · 2048 in / 128 out
latency
41 ms
Latency P50
41 ms
Throughput
1,020 ms
Power
620 W
Gemma latency profile on NVIDIA H200.
Whisper Large v3
INT8 Streaming ASR · INT8 · INT8
streaming-asr · ONNX Runtime · Batch 4 · 10s chunk
throughput
34 x realtime
Latency P50
88 ms
Throughput
34 x realtime
Power
610 W
Streaming Whisper profile on NVIDIA H200.
YOLOv10m
INT8 1280x1280 · INT8 · INT8
object-detection · TensorRT · Batch 1 · 1280x1280
throughput
420 images/s
Latency P50
12 ms
Throughput
420 images/s
Power
640 W
High-resolution YOLO profile on NVIDIA H200.
ViT-B/16
FP16 384x384 · FP16
image-classification · ONNX Runtime · Batch 64 · 384x384
throughput
9,800 images/s
Latency P50
14 ms
Throughput
9,800 images/s
Power
620 W
High-resolution ViT profile on NVIDIA H200.
Llama 3.1 8B Instruct
INT4 Serve · FP8 · INT4
text-generation · TensorRT-LLM · Batch 32 · Seq 2048 · 2048 in / 256 out
throughput
8,300 tokens/s
Latency P50
71 ms
Throughput
8,300 tokens/s
Power
680 W
Dense Llama INT4 serving profile on NVIDIA H200.
Qwen2.5 7B Instruct
FP8 Low-Latency · FP8
text-generation · vLLM · Batch 1 · Seq 4096 · 4096 in / 256 out
latency
49 ms
Latency P50
49 ms
Throughput
890 ms
Power
640 W
Low-latency Qwen profile on NVIDIA H200.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 1 · Seq 1024 · 1024 in / 128 out
latency
38 ms
Latency P50
38 ms
Throughput
—
Power
640 W
Demo single-request latency row for UX validation.
MMS English ASR
INT8 ASR · INT8 · INT8
speech-recognition · ONNX Runtime · Batch 8 · 15s clip
throughput
34.1 x realtime
Latency P50
—
Throughput
34.1 x realtime
Power
500 W
Demo compact speech row for category preview.
Whisper Large v3
FP16 ASR · FP16
speech-recognition · PyTorch · Batch 4 · 30s clip
throughput
27.5 x realtime
Latency P50
—
Throughput
27.5 x realtime
Power
510 W
Demo offline ASR row for category preview.
ViT-B/16
INT8 224x224 · INT8 · INT8
image-classification · ONNX Runtime · Batch 128 · 224x224
throughput
8,800 images/s
Latency P50
8 ms
Throughput
8,800 images/s
Power
520 W
Demo classification row for public category preview.
YOLOv10m
FP16 640x640 · FP16
object-detection · ONNX Runtime · Batch 32 · 640x640
throughput
6,200 images/s
Latency P50
12 ms
Throughput
6,200 images/s
Power
540 W
Demo object detection row for public category preview.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 8 · Seq 4096 · 4096 in / 512 out
throughput
6,100 tokens/s
Latency P50
52 ms
Throughput
6,100 tokens/s
Power
690 W
Demo composite throughput row used to validate public benchmark cards.
Qwen2.5 7B Instruct
BF16 Serve · BF16
text-generation · vLLM · Batch 16 · Seq 2048 · 2048 in / 256 out
throughput
6,900 tokens/s
Latency P50
60 ms
Throughput
6,900 tokens/s
Power
670 W
Demo multi-tenant throughput row for UX validation.
Specifications
ManufacturerNVIDIA
CategoryData Center GPU
ArchitectureHopper
Process Node4N
Form FactorSXM
CoolingLiquid or air
VRAM141 GB
VRAM Type—
Interconnect Bandwidth900 GB/s
Tensor / Matrix Cores528
Supported PrecisionsFP8, BF16, FP16, INT8
TDP700 W
FP16 (Dense)989 TFLOPS
FP32 (Dense)67 TFLOPS
Release Date2024-03-01
Price (USD)$32,000