activity
20212025
most citedCredit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

5 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

Yun Qu, Yuhang Jiang, Boyuan Wang +4

Reinforcement learning (RL) often encounters delayed and sparse feedback in real-world applications, even with only episodic rewards. Previous approaches have made some progress in…

cs.LG20221 cited

Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning

Zihan Zhang, Yuhang Jiang, Yuan Zhou +1

In this paper, we study the episodic reinforcement learning (RL) problem modeled by finite-horizon Markov Decision Processes (MDPs) with constraint on the number of batches. The mu…

cs.LG20211 cited

Reducing Conservativeness Oriented Offline Reinforcement Learning

Hongchang Zhang, Jianzhun Shao, Yuhang Jiang +2

In offline reinforcement learning, a policy learns to maximize cumulative rewards with a fixed collection of data. Towards conservative strategy, current methods choose to regulari…

cs.LG20215 cited

Credit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

Jianzhun Shao, Hongchang Zhang, Yuhang Jiang +2

Reward decomposition is a critical problem in centralized training with decentralized execution~(CTDE) paradigm for multi-agent reinforcement learning. To take full advantage of gl…

cs.CV2021

PFRL: Pose-Free Reinforcement Learning for 6D Pose Estimation

Jianzhun Shao, Yuhang Jiang, Gu Wang +2

6D pose estimation from a single RGB image is a challenging and vital task in computer vision. The current mainstream deep model methods resort to 2D images annotated with real-wor…