d-Matrix Corsair Inference Server
Eight Corsair cards in one server, rated by d-Matrix at 19.2 PFLOPS of MXINT8 and 76.8 PFLOPS of MXINT4, both labelled dense on its own page. Memory is two-tier and unusual: 16 GB of on-die performance memory running at 1,200 TB/s, backed by up to 2 TB of capacity memory. There is no HBM anywhere in the machine. d-Matrix quotes 60,000 tokens per second at 1 ms per token for Llama3 8B on a single server. It publishes no power figure for the card, the server or the rack, so total system power is left blank rather than estimated. Note that the 1,200 TB/s belongs to the on-die SRAM and is deliberately not recorded as memory bandwidth, which on this site means DRAM bandwidth; d-Matrix publishes no bandwidth for the capacity tier.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 19.200 PFLOPS | 2400.0 TFLOPS | — | |
| INT4 | 76.800 PFLOPS | 9600.0 TFLOPS | — |
All figures are dense. Vendors commonly headline the number, which is twice the dense one.
Powered by d-Matrix Corsair
This system utilizes 8 × d-Matrix Corsair GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
256 GB
Process Node
—
Architecture
Corsair
Related Systems
Documentation & Resources
Official Datasheet
d-Matrix Corsair Inference Server technical specifications
Corsair, Inference Server and Inference Rack specifications (d-Matrix)
d-Matrix • Latest version
Typical Use Cases
The d-Matrix Corsair Inference Server runs 8× Corsair GPUs over PCIe-based scale-up; 512 GB/s card-to-card DMX Bridge, delivering 19.2 PFLOPS INT8 dense.