activity
20182024
most citedCalibration, Entropy Rates, and Memory in Language Models

13 citations · 19 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2024

Provable Length Generalization in Sequence Prediction via Spectral Filtering

Annie Marsden, Evan Dogariu, Naman Agarwal +3

We consider the problem of length generalization in sequence prediction. We define a new metric of performance in this setting -- the Asymmetric-Regret -- which measures regret aga…

cs.LG2024

FutureFill: Fast Generation from Convolutional Sequence Models

Naman Agarwal, Xinyi Chen, Evan Dogariu +6

We address the challenge of efficient auto-regressive generation in sequence prediction models by introducing FutureFill, a general-purpose fast generation method for any sequence…

cs.LG2020

Black-Box Control for Linear Dynamical Systems

Xinyi Chen, Elad Hazan

We consider the problem of controlling an unknown linear time-invariant dynamical system from a single chain of black-box interactions, with no access to resets or offline simulati…

cs.LG20204 cited

Online Agnostic Boosting via Regret Minimization

Nataly Brukhim, Xinyi Chen, Elad Hazan +1

Boosting is a widely used machine learning approach based on the idea of aggregating weak learning rules. While in statistical learning numerous boosting methods exist both in the…

cs.LG2019

Extreme Tensoring for Low-Memory Preconditioning

Xinyi Chen, Naman Agarwal, Elad Hazan +2

State-of-the-art models are now trained with billions of parameters, reaching hardware limits in terms of memory consumption. This has created a recent demand for memory-efficient…

cs.LG2018

Efficient Full-Matrix Adaptive Regularization

Naman Agarwal, Brian Bullins, Xinyi Chen +4

Adaptive regularization methods pre-multiply a descent direction by a preconditioning matrix. Due to the large number of parameters of machine learning problems, full-matrix precon…