2 papers
cs.LG2026
Low-Rank Attention Residuals
Jonathan Su
Attention Residuals replace the fixed residual sum with depthwise attention over previous sub-layer outputs in large language models (LLMs), but use each output as both a full-dime…
cs.CL2026
Attention Projection Mixing with Exogenous Anchors
Jonathan Su
Cross-layer reuse of early attention projections can improve optimization and data efficiency, but it creates a structural conflict: the first layer must simultaneously act as a st…