3 papers
cs.LG2026
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
Xixi Wu, Qianguo Sun, Ruiyang Zhang +4
Reinforcement Learning (RL) is essential for evolving Large Language Models (LLMs) into autonomous agents capable of long-horizon planning, yet a practical recipe for scaling RL in…
cs.LG2025
Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
Rui Wang, Qianguo Sun, Chao Song +4
DPO (Direct Preference Optimization) has become a widely used offline preference optimization algorithm due to its simplicity and training stability. However, DPO is prone to overf…
cs.SD2025
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
Rui Wang, Qianguo Sun, Tianrong Chen +3
The emergence of multi-codebook neutral audio codecs such as Residual Vector Quantization (RVQ) and Group Vector Quantization (GVQ) has significantly advanced Large-Language-Model…