3 papers
cs.AI2026
From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents
Xingyu Su, Abhishek Kumar, Qing Ping +5
On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the sa…
cs.RO2026
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
Zeyun Deng, Yuzhe Lu, Yawei Wang +6
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes…
cs.CL2025
Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
Yuchen Zhuang, Jingfeng Yang, Haoming Jiang +16
Due to the scarcity of agent-oriented pre-training data, LLM-based autonomous agents typically rely on complex prompting or extensive fine-tuning, which often fails to introduce ne…