This part has no DRAM of any kind. Its 230 MB of SRAM is the whole memory system, which is why vram_gb is empty and memory_type reads On-die SRAM, and the 80 TB/s sits in the effective bandwidth column because it is an on-die figure and is not comparable with an HBM or GDDR number. The power figures are the CHIP'S, from Groq's GroqChip brief: Max 300 W, TDP 215 W, Average 185 W. The GroqCard that carries it is rated higher, Max 375 W, TDP 275 W and Typical 240 W, the difference being the board, its regulators and its RealScale PHYs. The two briefs also differ on link count by design: the chip integrates 16 RealScale interconnects and the card exposes up to 11 of them, with SKUs at 11, 9 and 0. INT8 is four times FP16 on this chip rather than twice, and that is correct rather than a transcription error: the matrix unit is an integer array whose native types are INT8, INT16 and INT32, and floating point is reached through Groq's TruePoint decomposition, so an FP16 multiply costs four integer operations. 750 divided by 4 is 187.5, which is the 188 Groq prints. Groq states that the design has no sparsity optimisation, so sparsity_support is 0 rather than NULL and neither figure has a sparse partner. This is not the same part as the NVIDIA Groq 3 LPU in this catalogue, which is the third-generation LP30 announced at GTC 2026 after NVIDIA acquired Groq.
Groq logo

Groq GroqChip 1

Type: LPUArchitecture: Tensor Streaming ProcessorForm factor: PCIeReleased: 2019Process: 14nmSpec confidence: Official
TDP
215 W

Overview

The GroqChip 1 is the first Tensor Streaming Processor, announced in November 2019 and sold from 2022 as the GroqCard accelerator. It is a 25 by 29 millimetre 14nm die of 26.8 billion transistors running at a fixed 900 MHz, rated at 750 TOPS of INT8 and 188 TFLOPS of FP16, and it has no external memory at all: 230 MB of globally shared SRAM at up to 80 TB/s is the entire memory system. Everything about it follows from that. There is no cache, no branch prediction and no hardware scheduler, so the compiler places every instruction on a known cycle and the same program takes exactly the same time on every run, which is what Groq sells rather than peak arithmetic. A model larger than 230 MB has to be split across chips, so the die spends its perimeter on sixteen RealScale links that connect cards directly without a switch. Against a GPU of the same era the memory is three orders of magnitude smaller and the latency is the product. This is first-generation Groq silicon and not the NVIDIA-branded Groq 3 LPU, which is three generations and one acquisition later.

Performance

Peak theoretical throughput by precision type

PrecisionPeak
FP64
No verified data available
FP32
32-bit floating point
No verified data available
TF32
No verified data available
BF16
No verified data available
FP16
16-bit floating point
188TFLOPS
FP8
No verified data available
FP6
No verified data available
FP4
No verified data available
INT8
8-bit integer
750TOPS

Every figure here is dense. The vendor documents no structured (2:4) sparsity mode for this part, so there is no second number to quote.

The GroqChip 1 in the GPU landscape

Peak FP16 TFLOPS (dense) against TDP, single-GPU parts tracked by Flopper

06001,2001,8002,4003,0000 W250 W500 W750 W1000 W1250 W1500 WInstinct MI355XGroqChip 1

Higher and further left is better: more half-precision throughput for less power.

Specifications

Architecture

Tensor Streaming Processor

Form Factor

PCIe

Launch Year

2019

Process Node

14nm

Memory

No verified data available On-die SRAM

Bandwidth

No verified data available

TDP

215 W

Max power

300 W

Transistors

26.8 billion

Interconnect

RealScale

Spec Confidence

Official

Full Specifications

Compute Engine
Base Clock 900 MHz
Chip Design
Transistors 26.8 billion
Die Size 725 mm²
Process Node 14nm
Memory
Memory No verified data available
Memory Type On-die SRAM
Bandwidth No verified data available
Interface Width No verified data available
Effective Bandwidth 80 TB/s
On-Die SRAM 230 MB
Interconnect & I/O
GPU-to-GPU RealScale
Interconnect Links 16
PCIe 4.0 x16
Power & Thermal
TDP 215 W
Max power 300 W
Enterprise Features
ECC Memory Yes
Sparsity No
General
Form Factor PCIe
Architecture Tensor Streaming Processor
Process Node 14nm
Launch Year 2019

Datasheet & Resources

Flopper Datasheet

Groq GroqChip 1 specifications, generated from our database. Printable.

GroqChip Processor Product Brief v1.5 (Groq)

Groq · Latest version

View

Data Provenance

Every figure traced to a source

Primary Source

Publisher
Groq
Published
No verified data available

Data Quality

Spec confidence
Official
Clock basis
Boost
Core precisions with figures
2 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Frequently Asked Questions

How many TFLOPS does the Groq GroqChip 1 have?

The Groq GroqChip 1 delivers 188 TFLOPS FP16 at peak. Flopper does not currently have verified FP32 and FP8 throughput figures for it.

What is the power consumption of the Groq GroqChip 1?

The Groq GroqChip 1 has a TDP (Thermal Design Power) rating of 215 watts.

How much memory does the Groq GroqChip 1 have?

Flopper does not currently have verified memory capacity or bandwidth figures for the Groq GroqChip 1.

What architecture is the Groq GroqChip 1 based on?

The Groq GroqChip 1 is based on the Tensor Streaming Processor architecture, launched in 2019.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs

© 2026 Flopper.io - Compare the GPUs Powering AI