d-Matrix logo

d-Matrix Corsair SquadRack

rack
64× Corsair
Scale-out over PCIe or Ethernet; 512 GB/s card-to-card DMX Bridge within each server
INT8
154
POPS
INT4
614
POPS
Power
—
kW Total
Memory
16400
GB Total

d-Matrix's rack-scale configuration, marketed as SquadRack and specified on its product page under the heading "Inference Rack": eight Corsair Inference Servers holding 64 cards in total. It carries 128 GB of on-die performance memory at 9.6 PB/s and up to 16.4 TB of capacity memory, with no HBM at all, and d-Matrix frames the two tiers as separate operating modes — up to 100B-parameter models held in performance mode, and 1T+ frontier models in capacity mode. Scale-out is over PCIe or Ethernet. d-Matrix quotes 30,000 tokens per second at 2 ms per token for Llama3 70B on a single rack. Note that d-Matrix publishes no compute figure for the rack: the 153.6 PFLOPS MXINT8 and 614.4 PFLOPS MXINT4 here are eight times its published server figures, which is safe because its own rack memory figure of 128 GB is likewise exactly eight times the server's 16 GB, and both reconcile against the 64 cards. No power figure is published for any part of the product line.

We will point you at suppliers who have it. Free, and no signup.

INT8
154
POPS
INT4
614
POPS

System Details

GPU Configuration

GPU Count: 64 GPUs
Architecture: Corsair
Interconnect: Scale-out over PCIe or Ethernet; 512 GB/s card-to-card DMX Bridge within each server

System Specifications

Form Factor: rack
Total Power: —
Total Memory: 16.4 TB
On-die SRAM: 128 GB
Memory Bandwidth: —

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
154 POPS 2,400 TOPS —
INT4 614 POPS 9,600 TOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Published throughput

Model Tokens/sec Basis Conditions
Llama3 70B 30,000 Whole system 2 ms/token

Tokens per second is a benchmark result, not a hardware specification. It depends on the model, the precision and the latency target it was measured under, so figures are only comparable when every condition matches. Vendor claims unless stated otherwise.

Powered by d-Matrix Corsair

This system utilizes 64 × d-Matrix Corsair GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

256 GB

Process Node

—

Architecture

Corsair

Documentation & Resources

Flopper Spec Sheet

d-Matrix Corsair SquadRack specifications, generated from our database. Printable.

Vendor product page

d-Matrix Corsair SquadRack documentation

View ↗

Corsair, Inference Server and Inference Rack specifications (d-Matrix)

d-Matrix • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The d-Matrix Corsair SquadRack runs 64× Corsair GPUs over Scale-out over PCIe or Ethernet; 512 GB/s card-to-card DMX Bridge within each server, delivering 154 POPS INT8 dense.

© 2026 Flopper.io - Compare the hardware powering AI