12 papers
Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation
Yunhao Zhao, Zhenyang Ni, Haoyang Chen +2
Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length…
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation
Jinhang Xu, Qiyuan Zhu, Yujun Wu +11
LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automation for whom? Researchers ope…
CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL
Zhenyang Ni, Yijiang Li, Ruochen Jiao +7
Video generation models trained on heterogeneous data with likelihood-surrogate objectives can produce visually plausible rollouts that violate physical constraints in embodied man…
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
Bowen Ding, Yuhan Chen, Jiayang Lyv +9
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ fundamentally in their reliance on expert t…
BehaviorGuard: Online Backdoor Defense for Deep Reinforcement Learning
Yinbo Yu, Xueyu Yin, Jiadai Wang +4
Backdoor attacks pose a serious threat to deep reinforcement learning (DRL). Current defenses typically rely on reward anomalies to reverse-engineer triggers and model finetuning t…
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
Simon Sinong Zhan, Qingyuan Wu, Philip Wang +4
Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present…