collaborators

7 papers

cs.CL2026

Learning to Adapt SFT Data for Better Reasoning Generalization

Lisong Sun, Li Wang, Chen Zhang +4

Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Among post-training paradigms,…

cs.LG2026

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

Li Wang, Xiaodong Lu, Xiaohan Wang +5

Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…

cs.DC2026

TSFLora: Token-Compressed Split Fine-Tuning for Wireless Edge Networks

Xianke Qiang, Zheng Chang, Li Wang +1

Adapting large AI models (LAMs) to personalized edge data is challenging because wireless devices have limited memory, computation, and uplink capacity. Federated fine-tuning prese…

cs.CL2026

TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression

Li Wang, Yandong Wang, Xin Yu +3

The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model…

cs.CL2025

Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning

Li Wang, Changhao Zhang, Zengqi Xiu +4

Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…

cs.LG2025

An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms

Li Wang, Sudun, Xingjian Zhang +2

Batch Normalization (BN) has played a pivotal role in the success of deep learning by improving training stability, mitigating overfitting, and enabling more effective optimization…