model compression 1multi-teacher aggregation 1policy optimization 1reasoning models 1reinforcement learning 1teacher-student distillation 1
From the 1 of 12 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TCPO: Turn-Level Credit Policy Optimization
Sicong Liao, Zhi Chen, Yaohua Tang
Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and itera…
cs.AI2026
LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning
Yubin Wu, Zicheng Cai, Liping Ning +4
Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device agents are constrained by l…