3 papers
cs.RO2026
The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping
Qi Luo, Shuaijun Liu, Hao Zhao +5
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according…
cs.LG2026
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Yifan Ding, Xincheng Wei, Yoshua Y. Li +7
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stron…
cs.LG2026
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Yunjie Chen, Xiaoxin Chen, Fang Wang
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models…