8 papers
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
Shi Fu, Yingjie Wang, Shengchao Hu +2
Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the cor…
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
Guozheng Ma, Lu Li, Zilin Wang +4
Scaling neural networks has driven breakthrough advances in machine learning, yet this paradigm fails in deep reinforcement learning (DRL), where larger models often degrade perfor…
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
Jifeng Hu, Sili Huang, Zhejian Yang +6
Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-…
Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
Ziqing Fan, Siyuan Du, Shengchao Hu +5
Selecting high-quality pre-training data for large language models (LLMs) is crucial for enhancing their overall performance under limited computation budget, improving both traini…
Squeeze Out Tokens from Sample for Finer-Grained Data Governance
Weixiong Lin, Chen Ju, Haicheng Wang +8
Widely observed data scaling laws, in which error falls off as a power of the training size, demonstrate the diminishing returns of unselective data expansion. Hence, data governan…
Continual Task Learning through Adaptive Policy Self-Composition
Shengchao Hu, Yuhang Zhou, Ziqing Fan +4
Training a generalizable agent to continually learn a sequence of tasks from offline trajectories is a natural requirement for long-lived agents, yet remains a significant challeng…