Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
Mutian He, Philip N. Garner
Linear-attention models that compress the entire input sequence into a fixed-size recurrent state offer an efficient alternative to Transformers, but their finite memory induces fo…
cs.CL2024
Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
Mutian He, Philip N. Garner
Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often u…
cs.CL2024
Exploring neural oscillations during speech perception via surrogate gradient spiking neural networks
Alexandre Bittar, Philip N. Garner
Understanding cognitive processes in the brain demands sophisticated models capable of replicating neural dynamics at large scales. We present a physiologically inspired speech rec…