3 papers
cs.AI2026
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
Zhexin Hu, Li Wang, Xiaohan Wang +4
Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression methods may discard task-critical…
cs.RO2026
RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations
Shuoqin Zhang, Yixin Xiong, Xiru Gao +4
Human-in-the-loop reinforcement learning systems achieve near-perfect success on the workstation where they are trained, but collapse when the same robot is moved to a workstation…
cs.AI2026
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
Zhenlin Wei, Pu Jian, Yingzhuo Deng +6
The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO…