8 papers
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
Qihuang Zhong, Liang Ding, Juhua Liu +3
Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data without external supervision. H…
Factorize to Generalize: Retrieval-Guided Invariant-Dynamic Decomposition for Time Series Forecasting
Jinjin Chi, Lei Feng, Lulu Zhang +6
Time series foundation models (TSFMs) have recently achieved strong zero-shot forecasting performance through large-scale pretraining and retrieval-augmented prediction. However, o…
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
Weixiao Zhan, Yongcheng Jing, Leszek Rutkowski +1
Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also underpins model leakage risks. Our…
Resource-constrained Amazons chess decision framework integrating large language models and graph attention
Tianhao Qian, Zhuoxuan Li, Jinde Cao +2
Artificial intelligence has advanced significantly through the development of intelligent game-playing systems, providing rigorous testbeds for decision-making, strategic planning,…
Alternating Gradient Flow Utility: A Unified Metric for Structural Pruning and Dynamic Routing in Deep Networks
Tianhao Qian, Zhuoxuan Li, Jinde Cao +2
Efficient deep learning traditionally relies on static heuristics like weight magnitude or activation awareness (e.g., Wanda, RIA). While successful in unstructured settings, we ob…
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
Guozheng Ma, Lu Li, Zilin Wang +4
Scaling neural networks has driven breakthrough advances in machine learning, yet this paradigm fails in deep reinforcement learning (DRL), where larger models often degrade perfor…