4 papers
TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression
Li Wang, Yandong Wang, Xin Yu +3
The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model…
Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
Li Wang, Changhao Zhang, Zengqi Xiu +4
Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
Li Wang, Sudun, Xingjian Zhang +2
Batch Normalization (BN) has played a pivotal role in the success of deep learning by improving training stability, mitigating overfitting, and enabling more effective optimization…
Embedded Mean Field Reinforcement Learning for Perimeter-defense Game
Li Wang, Xin Yu, Xuxin Lv +2
With the rapid advancement of unmanned aerial vehicles (UAVs) and missile technologies, perimeter-defense game between attackers and defenders for the protection of critical region…