4 papers
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
Keqin Peng, Yuanxin Ouyang, Xuebo Liu +4
Reinforcement Learning with Verifiable Rewards (RLVR) can elicit strong multi-step reasoning, yet it often encourages overly verbose traces. Moreover, naive length penalties in gro…
Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation
Yu Zhong, Zihao Zhang, Rui Zhang +9
Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of har…
Learning Guarantee of Reward Modeling Using Deep Neural Networks
Yuanhang Luo, Yeheng Ge, Ruijian Han +1
In this work, we study the learning theory of reward modeling with pairwise comparison data using deep neural networks. We establish a novel non-asymptotic regret bound for deep re…
A Survey on Large Language Model-based Agents for Statistics and Data Science
Maojun Sun, Ruijian Han, Binyan Jiang +4
In recent years, data science agents powered by Large Language Models (LLMs), known as "data agents," have shown significant potential to transform the traditional data analysis pa…