30 citations · 63 across the 9 of their papers we have counts for
8 papers · 1 filter
Better-than-KL PAC-Bayes Bounds
Ilja Kuzborskij, Kwang-Sung Jun, Yulian Wu +2
Let be a sequence of random elements, where is a fixed scalar function, are independent random variables (data), and i…
Towards Training Without Depth Limits: Batch Normalization Without Gradient Explosion
Alexandru Meterez, Amir Joudaki, Francesco Orabona +3
Normalization layers are one of the key building blocks for deep neural networks. Several theoretical studies have shown that batch normalization improves the signal propagation, b…
Normalized Gradients for All
Francesco Orabona
In this short note, I show how to adapt to Hölder smoothness using normalized gradients in a black-box way. Moreover, the bound will depend on a novel notion of local Hölder smooth…
Implicit Interpretation of Importance Weight Aware Updates
Keyi Chen, Francesco Orabona
Due to its speed and simplicity, subgradient descent is one of the most used optimization algorithms in convex machine learning algorithms. However, tuning its learning rate is pro…
Generalized Implicit Follow-The-Regularized-Leader
Keyi Chen, Francesco Orabona
We propose a new class of online learning algorithms, generalized implicit Follow-The-Regularized-Leader (FTRL), that expands the scope of FTRL framework. Generalized implicit FTRL…
Tighter PAC-Bayes Bounds Through Coin-Betting
Kyoungseok Jang, Kwang-Sung Jun, Ilja Kuzborskij +1
We consider the problem of estimating the mean of a sequence of random elements where is a fixed scalar function, ar…