29 citations · 45 across the 6 of their papers we have counts for
6 papers · 1 filter
Sub-Linear Memory: How to Make Performers SLiM
Valerii Likhosherstov, Krzysztof Choromanski, Jared Davis +2
The Transformer architecture has revolutionized deep learning on sequential data, becoming ubiquitous in state-of-the-art solutions for a wide variety of applications. Yet vanilla…
An Ode to an ODE
Krzysztof Choromanski, Jared Quincy Davis, Valerii Likhosherstov +6
We present a new paradigm for Neural ODE algorithms, called ODEtoODE, where time-dependent parameters of the main flow evolve according to a matrix flow on the orthogonal group O(d…
Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan +8
Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dep…
Time Dependence in Non-Autonomous Neural ODEs
Jared Quincy Davis, Krzysztof Choromanski, Jake Varley +6
Neural Ordinary Differential Equations (ODEs) are elegant reinterpretations of deep networks where continuous time can replace the discrete notion of depth, ODE solvers perform for…
CWY Parametrization: a Solution for Parallelized Optimization of Orthogonal and Stiefel Matrices
Valerii Likhosherstov, Jared Davis, Krzysztof Choromanski +1
We introduce an efficient approach for optimization over orthogonal groups on highly parallel computation units such as GPUs or TPUs. As in earlier work, we parametrize an orthogon…
Stochastic Flows and Geometric Optimization on the Orthogonal Group
Krzysztof Choromanski, David Cheikhi, Jared Davis +12
We present a new class of stochastic, geometrically-driven optimization algorithms on the orthogonal group and naturally reductive homogeneous manifolds obtained from the ac…