5 papers · 1 filter
SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis
Zhuohang Fan, Beichen Zhang, Yuanfa Li +4
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-covera…
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes…
Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning
Wei Liu, Ting Long
Cross-domain offline reinforcement learning (CDRL) aims to improve policy learning in a target domain by leveraging data collected from a source domain. Existing works typically as…
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Xiongbin Wu, Zhihao Luo, Shanzhe Lei +7
Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception…
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
Zhizheng Jiang, Kang Zhao, Weikai Xu +5
Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise…