2 papers
cs.LG2025
Learning in Compact Spaces with Approximately Normalized Transformer
Jörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina +3
The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the re…
cs.LG2024
Improving Deep Learning Optimization through Constrained Parameter Regularization
Jörg K. H. Franke, Michael Hefenbrock, Gregor Koehler +1
Regularization is a critical component in deep learning. The most commonly used approach, weight decay, applies a constant penalty coefficient uniformly across all parameters. This…