29 citations · 45 across the 6 of their papers we have counts for
7 papers · 1 filter
Debiasing a First-order Heuristic for Approximate Bi-level Optimization
Valerii Likhosherstov, Xingyou Song, Krzysztof Choromanski +2
Approximate bi-level optimization (ABLO) consists of (outer-level) optimization problems, involving numerical (inner-level) optimization loops. While ABLO has many applications acr…
Sub-Linear Memory: How to Make Performers SLiM
Valerii Likhosherstov, Krzysztof Choromanski, Jared Davis +2
The Transformer architecture has revolutionized deep learning on sequential data, becoming ubiquitous in state-of-the-art solutions for a wide variety of applications. Yet vanilla…
An Ode to an ODE
Krzysztof Choromanski, Jared Quincy Davis, Valerii Likhosherstov +6
We present a new paradigm for Neural ODE algorithms, called ODEtoODE, where time-dependent parameters of the main flow evolve according to a matrix flow on the orthogonal group O(d…
Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan +8
Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dep…
Time Dependence in Non-Autonomous Neural ODEs
Jared Quincy Davis, Krzysztof Choromanski, Jake Varley +6
Neural Ordinary Differential Equations (ODEs) are elegant reinterpretations of deep networks where continuous time can replace the discrete notion of depth, ODE solvers perform for…
CWY Parametrization: a Solution for Parallelized Optimization of Orthogonal and Stiefel Matrices
Valerii Likhosherstov, Jared Davis, Krzysztof Choromanski +1
We introduce an efficient approach for optimization over orthogonal groups on highly parallel computation units such as GPUs or TPUs. As in earlier work, we parametrize an orthogon…