29 citations · 42 across the 7 of their papers we have counts for
10 papers
Chefs' Random Tables: Non-Trigonometric Random Features
Valerii Likhosherstov, Krzysztof Choromanski, Avinava Dubey +3
We introduce chefs' random tables (CRTs), a new class of non-trigonometric random features (RFs) to approximate Gaussian and softmax kernels. CRTs are an alternative to standard ra…
On the Expressive Power of Self-Attention Matrices
Valerii Likhosherstov, Krzysztof Choromanski, Adrian Weller
Transformer networks are able to capture patterns in data coming from many domains (text, images, videos, proteins, etc.) with little or no change to architecture components. We pe…
Debiasing a First-order Heuristic for Approximate Bi-level Optimization
Valerii Likhosherstov, Xingyou Song, Krzysztof Choromanski +2
Approximate bi-level optimization (ABLO) consists of (outer-level) optimization problems, involving numerical (inner-level) optimization loops. While ABLO has many applications acr…
Unlocking Pixels for Reinforcement Learning via Implicit Attention
Krzysztof Marcin Choromanski, Deepali Jain, Wenhao Yu +9
There has recently been significant interest in training reinforcement learning (RL) agents in vision-based environments. This poses many challenges, such as high dimensionality an…
Sub-Linear Memory: How to Make Performers SLiM
Valerii Likhosherstov, Krzysztof Choromanski, Jared Davis +2
The Transformer architecture has revolutionized deep learning on sequential data, becoming ubiquitous in state-of-the-art solutions for a wide variety of applications. Yet vanilla…
An Ode to an ODE
Krzysztof Choromanski, Jared Quincy Davis, Valerii Likhosherstov +6
We present a new paradigm for Neural ODE algorithms, called ODEtoODE, where time-dependent parameters of the main flow evolve according to a matrix flow on the orthogonal group O(d…