4 papers
L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation
Liwen Zhang, Dong Zhou, Guanghui Sun +4
Intra-vehicular robots in spacecraft help reduce astronaut workload and improve mission efficiency. Recent research focuses on using deep learning methods to achieve the acute cont…
Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization
Dongruo Zhou
We study the deterministic first-order oracle complexity of finding \(ε\)-stationary points in smooth nonconvex optimization when the objective satisfies higher-order smoothness a…
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
Ruhan Wang, Yu Yang, Zhishuai Liu +2
We study offline off-dynamics reinforcement learning (RL) to utilize data from an easily accessible source domain to enhance policy learning in a target domain with limited data. O…
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
Yue Yu, Qiwei Di, Quanquan Gu +1
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of- (BoN)…