1 paper
Weiguo Lu, Gangnan Yuan, Hong-kun Zhang +1
Neural networks in general, from MLPs and CNNs to attention-based Transformers, are constructed from layers of linear combinations followed by nonlinear operations such as ReLU, Si…