3 papers
cs.RO2026
FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation
Kexin Shi, Junyao Shi, Poorvi Hebbar +5
Real-world reinforcement learning for robotic manipulation remains challenging, and this difficulty is amplified for flow matching policies: applying policy gradient methods to the…
cs.LG2024
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model
Jing Zhang, Linjiajie Fang, Kexin Shi +2
``Distribution shift'' is the main obstacle to the success of offline reinforcement learning. A learning policy may take actions beyond the behavior policy's knowledge, referred to…
cs.IR2024
Enhanced Bayesian Personalized Ranking for Robust Hard Negative Sampling in Recommender Systems
Kexin Shi, Jing Zhang, Linjiajie Fang +2
In implicit collaborative filtering, hard negative mining techniques are developed to accelerate and enhance the recommendation model learning. However, the inadvertent selection o…