1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.LG2025
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
Da Xiao, Qingye Meng, Shengping Li +1
We propose MUltiway Dynamic Dense (MUDD) connections, a simple yet effective method to address the limitations of residual connections and enhance cross-layer information flow in T…
cs.LG2024★ 1 cited
Improving Transformers with Dynamically Composable Multi-Head Attention
Da Xiao, Qingye Meng, Shengping Li +1
Multi-Head Attention (MHA) is a key component of Transformer. In MHA, attention heads work independently, causing problems such as low-rank bottleneck of attention score matrices a…
cs.LG2018
Improving the Universality and Learnability of Neural Programmer-Interpreters with Combinator Abstraction
Da Xiao, Jo-Yu Liao, Xingyuan Yuan
To overcome the limitations of Neural Programmer-Interpreters (NPI) in its universality and learnability, we propose the incorporation of combinator abstraction into neural program…