53 citations · 53 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
Rui Bu, Haofeng Zhong, Wenzheng Chen +1
Large models based on the Transformer architecture are susceptible to extreme-token phenomena, such as attention sinks and value-state drains. These issues, which degrade model per…
cs.LG2019
Mutual Information Maximization in Graph Neural Networks
Xinhan Di, Pengqian Yu, Rui Bu +1
A variety of graph neural networks (GNNs) frameworks for representation learning on graphs have been recently developed. These frameworks rely on aggregation and iteration scheme t…