Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights
Qishuai Wen, Zhiyuan Huang, Xianghan Meng +2
The vanilla self-attention mechanism in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equ…
cs.LG2025
Neural Normalized Cut: A Differential and Generalizable Approach for Spectral Clustering
Wei He, Shangzhi Zhang, Chun-Guang Li +3
Spectral clustering, as a popular tool for data clustering, requires an eigen-decomposition step on a given affinity to obtain the spectral embedding. Nevertheless, such a step suf…