4 papers
Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights
Qishuai Wen, Zhiyuan Huang, Xianghan Meng +2
The vanilla self-attention mechanism in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equ…
Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs
Jin Bai, Huiyao Zhang, Qi Wen +4
The segmentation of thin linear structures is inherently topology allowbreak-critical, where minor local errors can sever long-range connectivity. While recent State-Space Models (…
Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
Qishuai Wen, Zhiyuan Huang, Chun-Guang Li
Attention mechanisms have achieved significant empirical success in multiple fields, but their underlying optimization objectives remain unclear yet. Moreover, the quadratic comple…
Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression Perspective
Qishuai Wen, Chun-Guang Li
State-of-the-art methods for Transformer-based semantic segmentation typically adopt Transformer decoders that are used to extract additional embeddings from image embeddings via c…