d-Matrix logo

d-Matrix Corsair Inference Server

rack
8× Corsair
PCIe-based scale-up; 512 GB/s card-to-card DMX Bridge
INT8
19,200
TOPS
INT4
76,800
TOPS
Power
—
kW Total
Memory
2000
GB Total

Eight Corsair cards in one server, rated by d-Matrix at 19.2 PFLOPS of MXINT8 and 76.8 PFLOPS of MXINT4, both labelled dense on its own page. Memory is two-tier and unusual: 16 GB of on-die performance memory running at 1,200 TB/s, backed by up to 2 TB of capacity memory. There is no HBM anywhere in the machine. d-Matrix quotes 60,000 tokens per second at 1 ms per token for Llama3 8B on a single server. It publishes no power figure for the card, the server or the rack, so total system power is left blank rather than estimated. Note that the 1,200 TB/s belongs to the on-die SRAM and is deliberately not recorded as memory bandwidth, which on this site means DRAM bandwidth; d-Matrix publishes no bandwidth for the capacity tier.

We will point you at suppliers who have it. Free, and no signup.

INT8
19,200
TOPS
INT4
76,800
TOPS

System Details

GPU Configuration

GPU Count: 8 GPUs
Architecture: Corsair
Interconnect: PCIe-based scale-up; 512 GB/s card-to-card DMX Bridge

System Specifications

Form Factor: rack
Total Power: —
Total Memory: 2.0 TB
On-die SRAM: 16 GB
Memory Bandwidth: —

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
19,200 TOPS 2,400 TOPS —
INT4 76,800 TOPS 9,600 TOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Published throughput

Model Tokens/sec Basis Conditions
Llama3 8B 60,000 Whole system 1 ms/token

Tokens per second is a benchmark result, not a hardware specification. It depends on the model, the precision and the latency target it was measured under, so figures are only comparable when every condition matches. Vendor claims unless stated otherwise.

Powered by d-Matrix Corsair

This system utilizes 8 × d-Matrix Corsair GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

256 GB

Process Node

—

Architecture

Corsair

Documentation & Resources

Flopper Spec Sheet

d-Matrix Corsair Inference Server specifications, generated from our database. Printable.

Vendor product page

d-Matrix Corsair Inference Server documentation

View ↗

Corsair, Inference Server and Inference Rack specifications (d-Matrix)

d-Matrix • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The d-Matrix Corsair Inference Server runs 8× Corsair GPUs over PCIe-based scale-up; 512 GB/s card-to-card DMX Bridge, delivering 19,200 TOPS INT8 dense.

© 2026 Flopper.io - Compare the hardware powering AI