1 citations · 1 across the 13 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TCPO: Turn-Level Credit Policy Optimization
Sicong Liao, Zhi Chen, Yaohua Tang
Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and itera…
cs.AI2026
LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning
Yubin Wu, Zicheng Cai, Liping Ning +4
Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device agents are constrained by l…