activity
20242026
collaborators

8 papers

cs.AI2026

Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models

Shi Fu, Yingjie Wang, Shengchao Hu +2

Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the cor…

cs.LG2025

Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning

Guozheng Ma, Lu Li, Zilin Wang +4

Scaling neural networks has driven breakthrough advances in machine learning, yet this paradigm fails in deep reinforcement learning (DRL), where larger models often degrade perfor…

cs.LG2025

Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

Jifeng Hu, Sili Huang, Zhejian Yang +6

Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-…

cs.LG2025

Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection

Ziqing Fan, Siyuan Du, Shengchao Hu +5

Selecting high-quality pre-training data for large language models (LLMs) is crucial for enhancing their overall performance under limited computation budget, improving both traini…

cs.LG2025

Squeeze Out Tokens from Sample for Finer-Grained Data Governance

Weixiong Lin, Chen Ju, Haicheng Wang +8

Widely observed data scaling laws, in which error falls off as a power of the training size, demonstrate the diminishing returns of unselective data expansion. Hence, data governan…

cs.LG2024

Continual Task Learning through Adaptive Policy Self-Composition

Shengchao Hu, Yuhang Zhou, Ziqing Fan +4

Training a generalizable agent to continually learn a sequence of tasks from offline trajectories is a natural requirement for long-lived agents, yet remains a significant challeng…