activity
20182022
most citedHierarchical Adaptive Contextual Bandits for Resource Constraint based Recommendation

8 citations · 25 across the 5 of their papers we have counts for

collaborators

8 papers

cs.LG20221 cited

Spatio-temporal Incentives Optimization for Ride-hailing Services with Offline Deep Reinforcement Learning

Yanqiu Wu, Qingyang Li, Zhiwei Qin

A fundamental question in any peer-to-peer ride-sharing system is how to, both effectively and efficiently, meet the request of passengers to balance the supply and demand in real…

cs.LG20225 cited

Reinforcement Learning in the Wild: Scalable RL Dispatching Algorithm Deployed in Ridehailing Marketplace

Soheil Sadeghi Eshkevari, Xiaocheng Tang, Zhiwei Qin +4

In this study, a real-time dispatching algorithm based on reinforcement learning is proposed and for the first time, is deployed in large scale. Current dispatching methods in ride…

cs.LG20214 cited

Value Function is All You Need: A Unified Learning Framework for Ride Hailing Platforms

Xiaocheng Tang, Fan Zhang, Zhiwei Qin +6

Large ride-hailing platforms, such as DiDi, Uber and Lyft, connect tens of thousands of vehicles in a city to millions of ride demands throughout the day, providing great promises…

cs.LG2021

Real-world Ride-hailing Vehicle Repositioning using Deep Reinforcement Learning

Yan Jiao, Xiaocheng Tang, Zhiwei Qin +4

We present a new practical framework based on deep reinforcement learning and decision-time planning for real-world vehicle repositioning on ride-hailing (a type of mobility-on-dem…

cs.LG2020

Bayesian Meta-reinforcement Learning for Traffic Signal Control

Yayi Zou, Zhiwei Qin

In recent years, there has been increasing amount of interest around meta reinforcement learning methods for traffic signal control, which have achieved better performance compared…

cs.LG20208 cited

Hierarchical Adaptive Contextual Bandits for Resource Constraint based Recommendation

Mengyue Yang, Qingyang Li, Zhiwei Qin +1

Contextual multi-armed bandit (MAB) achieves cutting-edge performance on a variety of problems. When it comes to real-world scenarios such as recommendation system and online adver…