activity
20182021
most citedIntelligent Electric Vehicle Charging Recommendation Based on Multi-Agent Reinforcement Learning

97 citations · 116 across the 7 of their papers we have counts for

collaborators

9 papers

cs.RO20211 cited

Reinforcement Learning with Evolutionary Trajectory Generator: A General Approach for Quadrupedal Locomotion

Haojie Shi, Bo Zhou, Hongsheng Zeng +6

Recently reinforcement learning (RL) has emerged as a promising approach for quadrupedal locomotion, which can save the manual effort in conventional approaches such as designing s…

cs.LG20211 cited

ADER:Adapting between Exploration and Robustness for Actor-Critic Methods

Bo Zhou, Kejiao Li, Hongsheng Zeng +2

Combining off-policy reinforcement learning methods with function approximators such as neural networks has been found to lead to overestimation of the value function and sub-optim…

cs.AI20213 cited

Action Set Based Policy Optimization for Safe Power Grid Management

Bo Zhou, Hongsheng Zeng, Yuecheng Liu +3

Maintaining the stability of the modern power grid is becoming increasingly difficult due to fluctuating power consumption, unstable power supply coming from renewable energies, an…

cs.LG202197 cited

Intelligent Electric Vehicle Charging Recommendation Based on Multi-Agent Reinforcement Learning

Weijia Zhang, Hao Liu, Fan Wang +4

Electric Vehicle (EV) has become a preferable choice in the modern transportation system due to its environmental and energy sustainability. However, in many large cities, EV drive…

cs.LG20196 cited

Efficient and Robust Reinforcement Learning with Uncertainty-based Value Expansion

Bo Zhou, Hongsheng Zeng, Fan Wang +2

By integrating dynamics models into model-free reinforcement learning (RL) methods, model-based value expansion (MVE) algorithms have shown a significant advantage in sample effici…

cs.IR2019

MBCAL: Sample Efficient and Variance Reduced Reinforcement Learning for Recommender Systems

Fan Wang, Xiaomin Fang, Lihang Liu +2

In recommender systems such as news feed stream, it is essential to optimize the long-term utilities in the continuous user-system interaction processes. Previous works have proved…