3 papers
cs.LG2026
Collaborative Parameter Learning: Mitigating Forgetting via Parameter-Level Gradient Analysis
Mutian Yang, Zisen Zhan, Yutong Chen +7
Catastrophic forgetting during knowledge injection impairs the ability of large language models to acquire new knowledge without overwriting previously mastered knowledge. Recent s…
cs.LG2026
How Out-of-Distribution Detection Learning Theory Enhances Transformer: Learnability and Reliability
Yijin Zhou, Yutang Ge, Wenyuan Xie +3
Transformers excel in natural language processing and computer vision tasks. However, they still face challenges in generalizing to Out-of-Distribution (OOD) datasets, i.e. data wh…
cs.LG2025
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
Hanwen Liu, Yixuan Ma, Shi Jin +1
Attention mechanism is a significant part of Transformer models. It helps extract features from embedded vectors by adding global information and its expressivity has been proved t…