| Precision | Peak (PFLOPS) | Per GPU (TFLOPS) | TFLOPS/W |
|---|---|---|---|
| FP32 | 0.3 | 32.0 | — |
| FP16 | 1.0 | 128.0 | — |
| INT8 | 2.0 | 256.0 | — |
| INT16 | 1.0 | 128.0 | — |
All figures are dense. Vendors commonly headline the sparse number, which is twice the dense one.
Chip-to-chip interconnect at 200 GB/s
The Kunlunxin R480-X8 is an eight-card accelerator group built from the 32 GB R200-8F, the second-generation XPU-R part. Kunlunxin rates it at 2,048 TOPS of INT8, 1,024 TOPS of INT16, 1,024 TFLOPS of FP16 and 256 TFLOPS of FP32, with 256 GB of GDDR6 across the eight cards, 4,096 GB/s of aggregate memory bandwidth and 200 GB/s of chip-to-chip interconnect, on a 7nm process with ECC throughout. Kunlunxin lists it for both inference and training. Every figure is published per-card times eight on Kunlunxin's own page, and each per-card term matches the R200-8F exactly, which is what identifies the eight cards as the 32 GB variant rather than the 16 GB R200. Kunlunxin labels none of these figures dense or sparse and publishes no sparse rate. It also publishes no power figure for the group; summing eight 160 W cards would count only the accelerators and exclude CPUs, fans and supply losses, so total system power is left blank rather than understated.