
d-Matrix Corsair Inference Server
Eight Corsair cards in one server, rated by d-Matrix at 19.2 PFLOPS of MXINT8 and 76.8 PFLOPS of MXINT4, both labelled dense on its own page. Memory is two-tier and unusual: 16 GB of on-die performance memory running at 1,200 TB/s, backed by up to 2 TB of capacity memory. There is no HBM anywhere in the machine. d-Matrix quotes 60,000 tokens per second at 1 ms per token for Llama3 8B on a single server. It publishes no power figure for the card, the server or the rack, so total system power is left blank rather than estimated. Note that the 1,200 TB/s belongs to the on-die SRAM and is deliberately not recorded as memory bandwidth, which on this site means DRAM bandwidth; d-Matrix publishes no bandwidth for the capacity tier.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 19,200 TOPS | 2,400 TOPS | — | |
| INT4 | 76,800 TOPS | 9,600 TOPS | — |
Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.
Published throughput
| Model | Tokens/sec | Basis | Conditions |
|---|---|---|---|
| Llama3 8B | 60,000 | Whole system | 1 ms/token |
Tokens per second is a benchmark result, not a hardware specification. It depends on the model, the precision and the latency target it was measured under, so figures are only comparable when every condition matches. Vendor claims unless stated otherwise.
Powered by d-Matrix Corsair
This system utilizes 8 × d-Matrix Corsair GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
256 GB
Process Node
—
Architecture
Corsair
Related Systems
Documentation & Resources
Flopper Spec Sheet
d-Matrix Corsair Inference Server specifications, generated from our database. Printable.
Vendor product page
d-Matrix Corsair Inference Server documentation
Corsair, Inference Server and Inference Rack specifications (d-Matrix)
d-Matrix • Latest version
Typical Use Cases
The d-Matrix Corsair Inference Server runs 8× Corsair GPUs over PCIe-based scale-up; 512 GB/s card-to-card DMX Bridge, delivering 19,200 TOPS INT8 dense.