2 papers
cs.LG2026
Fast Weight Attention for Continual Learning
Yifan Zhang, Steve Ta, Jasper Zhang +8
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule…
cs.LG2026
DeepLoop: Depth Scaling for Looped Transformers
Shuzhen Li, Yifan Zhang, Jiacheng Guo +2
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters.…