2 papers
cs.LG2026
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
Yihong Chen, Zhouchen Lin, Quanming Yao
Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pr…
cs.LG2025
ReAugment: Model Zoo-Guided RL for Few-Shot Time Series Augmentation and Forecasting
Haochen Yuan, Yutong Wang, Yihong Chen +2
Time series forecasting, particularly in few-shot learning scenarios, is challenging due to the limited availability of high-quality training data. To address this, we present a pi…