12 citations · 34 across the 16 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024★ 2 cited
TransformerFAM: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2
While Transformers have revolutionized deep learning, their quadratic attention complexity hinders their ability to process infinitely long inputs. We propose Feedback Attention Me…
eess.AS2024
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models
Tsendsuren Munkhdalai, Youzheng Chen, Khe Chai Sim +3
Parameter efficient adaptation methods have become a key mechanism to train large pre-trained models for downstream tasks. However, their per-task parameter overhead is considered…