[D] Difference between CUDA and Tensor Cores

3DHydroPrints@alien.top · 2 years ago

[D] Difference between CUDA and Tensor Cores

VirtualHat@alien.top · 2 years ago

This big difference with tensor cores is that they use a 16-bit float multiply combined with a 32-bit float accumulate. This makes them much more efficient in terms of transistors required… but not a swap in replacement for CUDA.

Libraries like Pytorch can do matrix multiply (MM) on both CUDA cores and Tensor cores (and CPU, too, if you like). Typically Tensor cores are ~1.5-2x faster (in theory they’re much faster, in practice we’re often memory bandwidth limited so it doesn’t matter). The current default in Pytorch is to perform MM on CUDA, and convolutions on Tensor cores. The reason being that MM sometimes requires extra precision, and in vision models, most of the work is in the convolutions anyway.

wen_mars@alien.top · 2 years ago

And recently tensor cores have started appearing with 8 bit float/int as well, which gives them a huge advantage in inference throughput. The memory bandwidth limitation can be mitigated by increasing the batch size.

Buddy77777@alien.top · 2 years ago

Why use 32-bit accumulator to accumulate 16-bit numbers?

wen_mars@alien.top · 2 years ago

If you multiply two 16-bit numbers the result can overflow the range that can be represented by 16 bits.