activity
20132021
most citedMasked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

29 citations · 45 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2021

Debiasing a First-order Heuristic for Approximate Bi-level Optimization

Valerii Likhosherstov, Xingyou Song, Krzysztof Choromanski +2

Approximate bi-level optimization (ABLO) consists of (outer-level) optimization problems, involving numerical (inner-level) optimization loops. While ABLO has many applications acr…

cs.LG20205 cited

Sub-Linear Memory: How to Make Performers SLiM

Valerii Likhosherstov, Krzysztof Choromanski, Jared Davis +2

The Transformer architecture has revolutionized deep learning on sequential data, becoming ubiquitous in state-of-the-art solutions for a wide variety of applications. Yet vanilla…

cs.LG2020

An Ode to an ODE

Krzysztof Choromanski, Jared Quincy Davis, Valerii Likhosherstov +6

We present a new paradigm for Neural ODE algorithms, called ODEtoODE, where time-dependent parameters of the main flow evolve according to a matrix flow on the orthogonal group O(d…

cs.LG202029 cited

Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Krzysztof Choromanski, Valerii Likhosherstov, David Dohan +8

Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dep…

cs.LG20207 cited

Time Dependence in Non-Autonomous Neural ODEs

Jared Quincy Davis, Krzysztof Choromanski, Jake Varley +6

Neural Ordinary Differential Equations (ODEs) are elegant reinterpretations of deep networks where continuous time can replace the discrete notion of depth, ODE solvers perform for…

cs.LG2020

CWY Parametrization: a Solution for Parallelized Optimization of Orthogonal and Stiefel Matrices

Valerii Likhosherstov, Jared Davis, Krzysztof Choromanski +1

We introduce an efficient approach for optimization over orthogonal groups on highly parallel computation units such as GPUs or TPUs. As in earlier work, we parametrize an orthogon…