64 citations · 99 across the 8 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.LG2023★ 5 cited
Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
Ido Amos, Jonathan Berant, Ankit Gupta
Modeling long-range dependencies across sequences is a longstanding goal in machine learning and has led to architectures, such as state space models, that dramatically outperform…
eess.AS2023
Diagonal State Space Augmented Transformers for Speech Recognition
George Saon, Ankit Gupta, Xiaodong Cui
We improve on the popular conformer architecture by replacing the depthwise temporal convolutions with diagonal state space (DSS) models. DSS is a recently introduced variant of li…