2 papers
cs.LG2026
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
Zhong Guan, Yongjian Guo, Haoran Sun +5
Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a c…
cs.RO2026
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
Jiyao Zhang, Zimu Han, Junhan Wang +7
Robotic imitation learning faces a fundamental trade-off between modeling long-horizon dependencies and enabling fine-grained closed-loop control. Existing fixed-frequency action c…