IntelAI Accelerator
Intel Gaudi 3
Based on Gaudi 3 architecture. 128GB memory with 900W TDP.
FP16 Performance
1835
TFLOPS
Memory Capacity
128
GB
Power Draw
900
Watts TDP
Performance per Dollar
0.08 TFLOPS / $
Based on list price
Subject to market fluctuation
Benchmark Results
Published results grouped by model variant and scenario definition
7 resultsOpen Category View
Mistral 7B Instruct v0.3
FP8 Dense Serve · FP8
text-generation · vLLM · Batch 16 · Seq 4096 · 4096 in / 256 out
throughput
5,020 tokens/s
Latency P50
69 ms
Throughput
5,020 tokens/s
Power
790 W
Dense Mistral serving profile on Intel Gaudi 3.
Gemma 2 9B Instruct
INT4 Chat Serve · BF16 · INT4
text-generation · TensorRT-LLM · Batch 4 · Seq 2048 · 2048 in / 128 out
latency
55 ms
Latency P50
55 ms
Throughput
820 ms
Power
770 W
Gemma latency profile on Intel Gaudi 3.
Llama 3.1 8B Instruct
INT4 Serve · FP8 · INT4
text-generation · TensorRT-LLM · Batch 32 · Seq 2048 · 2048 in / 256 out
throughput
6,900 tokens/s
Latency P50
85 ms
Throughput
6,900 tokens/s
Power
810 W
Dense Llama INT4 serving profile on Intel Gaudi 3.
Qwen2.5 7B Instruct
FP8 Low-Latency · FP8
text-generation · vLLM · Batch 1 · Seq 4096 · 4096 in / 256 out
latency
63 ms
Latency P50
63 ms
Throughput
690 ms
Power
780 W
Low-latency Qwen profile on Intel Gaudi 3.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 1 · Seq 1024 · 1024 in / 128 out
latency
55 ms
Latency P50
55 ms
Throughput
—
Power
830 W
Demo single-request latency row for UX validation.
Qwen2.5 7B Instruct
BF16 Serve · BF16
text-generation · vLLM · Batch 16 · Seq 2048 · 2048 in / 256 out
throughput
5,200 tokens/s
Latency P50
71 ms
Throughput
5,200 tokens/s
Power
845 W
Demo multi-tenant throughput row for UX validation.
Llama 3.1 8B Instruct
FP8 Serve · FP8 · FP8
text-generation · TensorRT-LLM · Batch 8 · Seq 4096 · 4096 in / 512 out
throughput
4,700 tokens/s
Latency P50
65 ms
Throughput
4,700 tokens/s
Power
860 W
Demo composite throughput row used to validate public benchmark cards.
Specifications
ManufacturerIntel
CategoryAI Accelerator
ArchitectureGaudi 3
Process Node5nm
Form FactorOAM
CoolingAir
VRAM128 GB
VRAM Type—
Interconnect Bandwidth512 GB/s
Tensor / Matrix Cores—
Supported PrecisionsBF16, FP16, FP8, INT8
TDP900 W
FP16 (Dense)1835 TFLOPS
FP32 (Dense)120 TFLOPS
Release Date2024-04-01
Price (USD)$22,000