1 citations · 1 across the 7 of their papers we have counts for
4 papers · 1 filter
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
Qibin Wang, Pu Zhao, Shaohan Huang +6
Test-time scaling (TTS) has gained widespread attention for enhancing LLM reasoning. Existing approaches such as Best-of-N and majority voting are limited as their performance depe…
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
Chenghua Huang, Lu Wang, Fangkai Yang +6
In this paper, we explore how directly pretraining a value model simplifies and stabilizes reinforcement learning from human feedback (RLHF). In reinforcement learning, value estim…
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
Jiani Zheng, Lu Wang, Fangkai Yang +7
Training Vision-Language Models (VLMs) for Graphical User Interfaces (GUI) agents via Reinforcement Learning (RL) faces critical challenges: environment-based RL requires costly in…
Token-level Proximal Policy Optimization for Query Generation
Yichen Ouyang, Lu Wang, Fangkai Yang +13
Query generation is a critical task for web search engines (e.g. Google, Bing) and recommendation systems. Recently, state-of-the-art query generation methods leverage Large Langua…