
d-Matrix Corsair SquadRack
d-Matrix's rack-scale configuration, marketed as SquadRack and specified on its product page under the heading "Inference Rack": eight Corsair Inference Servers holding 64 cards in total. It carries 128 GB of on-die performance memory at 9.6 PB/s and up to 16.4 TB of capacity memory, with no HBM at all, and d-Matrix frames the two tiers as separate operating modes — up to 100B-parameter models held in performance mode, and 1T+ frontier models in capacity mode. Scale-out is over PCIe or Ethernet. d-Matrix quotes 30,000 tokens per second at 2 ms per token for Llama3 70B on a single rack. Note that d-Matrix publishes no compute figure for the rack: the 153.6 PFLOPS MXINT8 and 614.4 PFLOPS MXINT4 here are eight times its published server figures, which is safe because its own rack memory figure of 128 GB is likewise exactly eight times the server's 16 GB, and both reconcile against the 64 cards. No power figure is published for any part of the product line.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 154 POPS | 2,400 TOPS | — | |
| INT4 | 614 POPS | 9,600 TOPS | — |
Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.
Published throughput
| Model | Tokens/sec | Basis | Conditions |
|---|---|---|---|
| Llama3 70B | 30,000 | Whole system | 2 ms/token |
Tokens per second is a benchmark result, not a hardware specification. It depends on the model, the precision and the latency target it was measured under, so figures are only comparable when every condition matches. Vendor claims unless stated otherwise.
Powered by d-Matrix Corsair
This system utilizes 64 × d-Matrix Corsair GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
256 GB
Process Node
—
Architecture
Corsair
Related Systems
Documentation & Resources
Flopper Spec Sheet
d-Matrix Corsair SquadRack specifications, generated from our database. Printable.
Vendor product page
d-Matrix Corsair SquadRack documentation
Corsair, Inference Server and Inference Rack specifications (d-Matrix)
d-Matrix • Latest version
Typical Use Cases
The d-Matrix Corsair SquadRack runs 64× Corsair GPUs over Scale-out over PCIe or Ethernet; 512 GB/s card-to-card DMX Bridge within each server, delivering 154 POPS INT8 dense.