vs 

NVIDIA Groq 3 LPU vs Groq GroqChip 1
Side-by-side specifications, performance, and rental pricing for datacenter AI workloads.
Groq 3 LPU: Groq LPU, 2026 GroqChip 1: Tensor Streaming Processor, 2019
NVIDIA Groq 3 LPU
Full specs →Architecture
Groq LPU
Year
2026

Groq GroqChip 1
Full specs →Architecture
Tensor Streaming Processor
Year
2019
TDP
215 W
Specifications
Groq 3 LPU
GroqChip 1

Architecture
Groq LPUTensor Streaming Processor
Launch Year
20262019
Form Factor
—PCIe
Memory
——
Memory Bandwidth
——
TDP
—215 W
Process Node
—14nm
Spec Groq 3 LPU
GroqChip 1

Architecture Groq LPUTensor Streaming Processor
Launch Year 20262019
Form Factor —PCIe
Memory ——
Memory Bandwidth ——
TDP —215 W
Process Node —14nm
Performance (TFLOPS)
Groq 3 LPU
GroqChip 1

FP64
No verified data available No verified data available
FP32
No verified data available No verified data available
TF32
No verified data available No verified data available
BF16
No verified data available No verified data available
FP16
No verified data available188 TFLOPS
FP8
1,200 TFLOPS No verified data available
FP6
No verified data available No verified data available
FP4
No verified data available No verified data available
INT8
No verified data available750 TOPS
Precision Groq 3 LPU
GroqChip 1

FP64 No verified data available No verified data available
FP32 No verified data available No verified data available
TF32 No verified data available No verified data available
BF16 No verified data available No verified data available
FP16 No verified data available188 TFLOPS
FP8 1,200 TFLOPS No verified data available
FP6 No verified data available No verified data available
FP4 No verified data available No verified data available
INT8 No verified data available750 TOPS
FLOPS by Precision
What actually differs
The Groq 3 LPU is the newer part: Groq LPU, launched in 2026, against the GroqChip 1's Tensor Streaming Processor from 2019. Newer architectures typically add lower-precision formats and better throughput per watt, so check the precision rows your workload actually uses.
For LLM training and inference, weight the FP8 and FP16 rows and memory capacity most heavily. For scientific and HPC workloads, the FP64 row is the one to read.
Get Comparison Updates
New GPUs added weekly. Be the first to see how they compare.