[D] What academic papers/textbooks do you recommend to learn about the ReLU activation function?

fouried96@alien.top · 1 year ago

[D] What academic papers/textbooks do you recommend to learn about the ReLU activation function?

d84-n1nj4@alien.top · 1 year ago

I believe it was first created in “Cognitron: A self organizing multilayered neural network”, but was not referred to as ReLU. It was popularized by “Deep Sparse Rectifier Neural Networks” and “Rectified Linear Units Improve Restricted Boltzmann Machines”.

In regard to deep learning and GPU use: It’s efficient compared to other activation functions because it consists of comparison and thresholding operations, and the derivative is just 1 when positive and 0 if not (for backpropagation). It’s effective because it adds non-linearity to layers of linear operations like the convolution.