8 papers
ANO: A Principled Approach to Robust Policy Optimization
Yiheng Zhang, Yiming Wang, Kaiyan Zhao +3
Proximal Policy Optimization (PPO) dominates reinforcement learning and LLM alignment but relies on a "hard clipping" mechanism that discards valuable gradients. Conversely, uncons…
Anon: Extrapolating Adaptivity Beyond SGD and Adam
Yiheng Zhang, Kaiyan Zhao, Shaowu Wu +5
Adaptive optimizers such as Adam have achieved great success in training large-scale models like large language models and diffusion models. However, they often generalize worse th…
EDT: Efficient and Effective Decision Transformer with Experience-Aware Sampling for Robotic Manipulation
Kaiyan Zhao, Borong Zhang, Yiming Wang +4
In reinforcement learning (RL) for robotic manipulation, the Decision Transformer (DT) has emerged as an effective framework for addressing long-horizon tasks. However, DT's perfor…
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma +9
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technolog…
When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens
Xuetao Li, Pinhan Fu, Wenke Huang +7
Downstream fine-tuning of vision-language-action (VLA) models enhances robotics, yet exposes the pipeline to backdoor risks. Attackers can pretrain VLAs on poisoned data to implant…
RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
Xuetao Li, Wenke Huang, Nengyuan Pan +7
Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate ext…