2 papers
cs.LG2026
HSD: Hybrid Hindsight Self-Distillation
Qiye Cai, Yichuan Ma, Linyang Li +7
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level…
cs.RO2026
Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control
Jun Chen, Erdent Bao, Wenlong Dong +7
Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step…