NVIDIA GeForce GTX 1070 Ti 8GB vs NVIDIA Tesla P40 24GB

The Tesla P40 delivers 47% more FP32 throughput than the GeForce GTX 1070 Ti (12 vs 8.2 TFLOPS dense). The Tesla P40 also carries 16 GB more memory (24 GB vs 8 GB).

GeForce GTX 1070 Ti: Pascal, 2017 Tesla P40: Pascal, 2016
FP32 dense lead
+47%
Tesla P40: 12 vs 8.2 TFLOPS
Memory
8 vs 24 GB
16 GB more for the Tesla P40
Bandwidth
256 GB/s vs 346 GB/s
Tesla P40 moves data faster
TDP
180 W vs 250 W
GeForce GTX 1070 Ti draws 70 W less

Specifications

GeForce GTX 1070 Ti
Tesla P40
Architecture
PascalPascal
Launch Year
20172016
Form Factor
PCIePCIe
Memory
8 GB24 GB
Memory Bandwidth
256 GB/s346 GB/s
TDP
180 W250 W
Process Node
16nm16nm

Performance (TFLOPS)

GeForce GTX 1070 Ti
Tesla P40
FP64
No verified data available No verified data available
FP32
8.2 TFlops 12.0 TFlops
TF32
No verified data available No verified data available
BF16
No verified data available No verified data available
FP16
No verified data available No verified data available
FP8
No verified data available No verified data available
FP6
No verified data available No verified data available
FP4
No verified data available No verified data available
INT8
No verified data available47.0 TFlops

FLOPS by Precision

What actually differs

The GeForce GTX 1070 Ti is the newer part: Pascal, launched in 2017, against the Tesla P40's Pascal from 2016. Newer architectures typically add lower-precision formats and better throughput per watt, so check the precision rows your workload actually uses.

Dense throughput for the GeForce GTX 1070 Ti against the Tesla P40: FP32 8.2 vs 12 TFLOPS. Sparse figures, where the vendor publishes them, appear under each dense number in the performance table.

Memory is 8 GB against 24 GB, fed at 256 GB/s versus 346 GB/s. More memory per GPU means larger models fit before you have to shard across cards, which often matters more than raw TFLOPS for inference.

Power budgets are 180 W for the GeForce GTX 1070 Ti and 250 W for the Tesla P40. At FP32 that works out to 0.05 against 0.05 TFLOPS per watt.

For LLM training and inference, weight the FP8 and FP16 rows and memory capacity most heavily. For scientific and HPC workloads, the FP64 row is the one to read.

Frequently asked questions

Is the GeForce GTX 1070 Ti faster than the Tesla P40?

At FP32 precision the Tesla P40 reaches 12 TFLOPS dense against 8.2 TFLOPS for the GeForce GTX 1070 Ti. The performance table on this page lists every published precision for both GPUs.

Which has more memory, the GeForce GTX 1070 Ti or the Tesla P40?

The Tesla P40 carries 24 GB of memory versus 8 GB for the GeForce GTX 1070 Ti. Memory bandwidth is 256 GB/s for the GeForce GTX 1070 Ti and 346 GB/s for the Tesla P40.

How much power do the GeForce GTX 1070 Ti and the Tesla P40 draw?

The GeForce GTX 1070 Ti is rated at 180 W TDP and the Tesla P40 at 250 W. On FP32 throughput per watt, the Tesla P40 is the more efficient part.

Get Comparison Updates

New GPUs added weekly. Be the first to see how they compare.