3 papers
cs.CL2026
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Hongcheng Wang, Yinuo Huang, Sukai Wang +2
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers f…
cs.RO2026
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
Ruiying Li, Yunlang Zhou, YuYao Zhu +15
Vision-Language-Action (VLA) systems have shown strong potential for language-driven robotic manipulation. However, scaling them to long-horizon tasks remains challenging. Existing…
cs.RO2026
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
Yi Liu, Sukai Wang, Dafeng Wei +10
General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challeng…