most citedMomentum Q-learning with Finite-Sample Convergence Guarantee

8 citations · 13 across the 7 of their papers we have counts for

collaborators

7 papers

cs.RO2020

Velocity Regulation of 3D Bipedal Walking Robots with Uncertain Dynamics Through Adaptive Neural Network Controller

Guillermo A. Castillo, Bowen Weng, Terrence C. Stewart +2

This paper presents a neural-network based adaptive feedback control structure to regulate the velocity of 3D bipedal robots under dynamics uncertainties. Existing Hybrid Zero Dyna…

cs.LG20208 cited

Momentum Q-learning with Finite-Sample Convergence Guarantee

Bowen Weng, Huaqing Xiong, Lin Zhao +2

Existing studies indicate that momentum ideas in conventional optimization can be used to improve the performance of Q-learning algorithms. However, the finite-sample analysis for…

cs.RO2020

Model Predictive Instantaneous Safety Metric for Evaluation of Automated Driving Systems

Bowen Weng, Sughosh J. Rao, Eeshan Deosthale +2

Vehicles with Automated Driving Systems (ADS) operate in a high-dimensional continuous system with multi-agent interactions. This continuous system features various types of traffi…

eess.SY20191 cited

Momentum-based Accelerated Q-learning

Bowen Weng, Lin Zhao, Huaqing Xiong +1

This paper studies accelerated algorithms for Q-learning. We propose an acceleration scheme by incorporating the historical iterates of the Q-function. The idea is conceptually ins…

cs.RO2019

Hybrid Zero Dynamics Inspired Feedback Control Policy Design for 3D Bipedal Locomotion using Reinforcement Learning

Guillermo A. Castillo, Bowen Weng, Wei Zhang +1

This paper presents a novel model-free reinforcement learning (RL) framework to design feedback control policies for 3D bipedal walking. Existing RL algorithms are often trained in…

cs.RO20194 cited

Reciprocal Collision Avoidance for General Nonlinear Agents using Reinforcement Learning

Hao Li, Bowen Weng, Abhishek Gupta +2

Finding feasible and collision-free paths for multiple nonlinear agents is challenging in the decentralized scenarios due to limited available information of other agents and compl…