13 citations · 13 across the 2 of their papers we have counts for
4 papers
No-Regret Prediction in Marginally Stable Systems
Udaya Ghai, Holden Lee, Karan Singh +2
We consider the problem of online prediction in a marginally stable linear dynamical system subject to bounded adversarial or (non-isotropic) stochastic perturbations. This poses t…
Calibration, Entropy Rates, and Memory in Language Models
Mark Braverman, Xinyi Chen, Sham M. Kakade +3
Building accurate language models that capture meaningful long-term dependencies is a core challenge in natural language processing. Towards this end, we present a calibration-base…
Extreme Tensoring for Low-Memory Preconditioning
Xinyi Chen, Naman Agarwal, Elad Hazan +2
State-of-the-art models are now trained with billions of parameters, reaching hardware limits in terms of memory consumption. This has created a recent demand for memory-efficient…
Efficient Full-Matrix Adaptive Regularization
Naman Agarwal, Brian Bullins, Xinyi Chen +4
Adaptive regularization methods pre-multiply a descent direction by a preconditioning matrix. Due to the large number of parameters of machine learning problems, full-matrix precon…