Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Towards Compressive and Scalable Recurrent Memory
Yunchong Song, Jushi Kai, Liming Lu +2
Transformers face a quadratic bottleneck in attention when scaling to long contexts. Recent approaches introduce recurrent memory to extend context beyond the current window, yet t…
cs.LG2025
FERD: Fairness-Enhanced Data-Free Robustness Distillation
Zhengxiao Li, Liming Lu, Xu Zheng +4
Data-Free Robustness Distillation (DFRD) aims to transfer the robustness from the teacher to the student without accessing the training data. While existing methods focus on overal…