NVIDIAData Center GPU

NVIDIA B200

Based on Blackwell architecture. 192GB memory with 1000W TDP.

FP16 Performance
1800
TFLOPS
Memory Capacity
192
GB
Power Draw
1000
Watts TDP
Performance per Dollar
0.04 TFLOPS / $
Based on list price
Subject to market fluctuation

Benchmark Results

Published results grouped by model variant and scenario definition
RT-DETR-L
FP16 960x960 · FP16
object-detection · TensorRT · Batch 8 · 960x960
throughput
1,210 images/s
Latency P50
11 ms
Throughput
1,210 images/s
Power
920 W
RT-DETR profile on NVIDIA B200.
Mistral 7B Instruct v0.3
FP8 Dense Serve · FP8
text-generation · vLLM · Batch 16 · Seq 4096 · 4096 in / 256 out
throughput
8,100 tokens/s
Latency P50
43 ms
Throughput
8,100 tokens/s
Power
910 W
Dense Mistral serving profile on NVIDIA B200.
Gemma 2 9B Instruct
INT4 Chat Serve · BF16 · INT4
text-generation · TensorRT-LLM · Batch 4 · Seq 2048 · 2048 in / 128 out
latency
31 ms
Latency P50
31 ms
Throughput
1,260 ms
Power
880 W
Gemma latency profile on NVIDIA B200.
Llama 3.1 8B Instruct
INT4 Serve · FP8 · INT4
text-generation · TensorRT-LLM · Batch 32 · Seq 2048 · 2048 in / 256 out
throughput
11,200 tokens/s
Latency P50
55 ms
Throughput
11,200 tokens/s
Power
930 W
Dense Llama INT4 serving profile on NVIDIA B200.
YOLOv10m
INT8 1280x1280 · INT8 · INT8
object-detection · TensorRT · Batch 1 · 1280x1280
throughput
560 images/s
Latency P50
9 ms
Throughput
560 images/s
Power
890 W
High-resolution YOLO profile on NVIDIA B200.
Qwen2.5 7B Instruct
FP8 Low-Latency · FP8
text-generation · vLLM · Batch 1 · Seq 4096 · 4096 in / 256 out
latency
36 ms
Latency P50
36 ms
Throughput
1,120 ms
Power
900 W
Low-latency Qwen profile on NVIDIA B200.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 8 · Seq 4096 · 4096 in / 512 out
throughput
9,200 tokens/s
Latency P50
39 ms
Throughput
9,200 tokens/s
Power
940 W
Demo composite throughput row used to validate public benchmark cards.
Qwen2.5 7B Instruct
BF16 Serve · BF16
text-generation · vLLM · Batch 16 · Seq 2048 · 2048 in / 256 out
throughput
10,400 tokens/s
Latency P50
44 ms
Throughput
10,400 tokens/s
Power
950 W
Demo multi-tenant throughput row for UX validation.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 1 · Seq 1024 · 1024 in / 128 out
latency
28 ms
Latency P50
28 ms
Throughput
Power
880 W
Demo single-request latency row for UX validation.

Specifications

ManufacturerNVIDIA
CategoryData Center GPU
ArchitectureBlackwell
Process NodeTSMC 4NP
Form FactorSXM
CoolingLiquid
VRAM192 GB
VRAM Type
Interconnect Bandwidth1800 GB/s
Tensor / Matrix Cores640
Supported PrecisionsFP8, BF16, FP16, INT8, INT4
TDP1000 W
FP16 (Dense)1800 TFLOPS
FP32 (Dense)90 TFLOPS
Release Date2025-01-01
Price (USD)$42,000