2 papers
cs.CV2025
ARFlow: Autoregressive Flow with Hybrid Linear Attention
Mude Hui, Rui-Jie Zhu, Songlin Yang +5
Flow models are effective at progressively generating realistic images, but they generally struggle to capture long-range dependencies during the generation process as they compres…
cs.CL2024
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
Yu Zhang, Songlin Yang, Ruijie Zhu +9
Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks comp…