3 papers
cs.AI2026
Multi-Head Attention Residuals
Cheng Luo, Zefan Cai, Junjie Hu
Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by lettin…
cs.LG2026
Delta Attention Residuals
Cheng Luo, Zefan Cai, Junjie Hu
Attention Residuals replace standard additive residual connections with learned softmax attention over previous layer outputs, enabling selective cross-layer routing. However, stan…
q-bio.GN2025
GenDMR: A dynamic multimodal role-swapping network for identifying risk gene phenotypes
Lina Qin, Cheng Zhu, Chuqi Zhou +8
Recent studies have shown that integrating multimodal data fusion techniques for imaging and genetic features is beneficial for the etiological analysis and predictive diagnosis of…