activity
20202024
most citedSelf6D: Self-Supervised Monocular 6D Object Pose Estimation

9 citations · 25 across the 9 of their papers we have counts for

collaborators

10 papers

cs.AI2024

Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration

Yun Qu, Boyuan Wang, Yuhang Jiang +7

With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing novelty, diversity, or uncertain…

cs.AI2024★ 1 cited

Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning Benchmarks

Yun Qu, Boyuan Wang, Jianzhun Shao +15

The advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected o…

cs.AI2024★ 1 cited

LLM-Empowered State Representation for Reinforcement Learning

Boyuan Wang, Yun Qu, Yuhang Jiang +4

Conventional state representations in reinforcement learning often omit critical task-related details, presenting a significant challenge for value networks in establishing accurat…

cs.AI2023★ 2 cited

Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement Learning

Jianzhun Shao, Yun Qu, Chen Chen +2

Offline multi-agent reinforcement learning is challenging due to the coupling effect of both distribution shift issue common in offline setting and the high dimension issue common…

cs.LG2021★ 3 cited

Wasserstein Unsupervised Reinforcement Learning

Shuncheng He, Yuhang Jiang, Hongchang Zhang +2

Unsupervised reinforcement learning aims to train agents to learn a handful of policies or skills in environments without external reward. These pre-trained policies can accelerate…

cs.LG2021★ 1 cited

Reducing Conservativeness Oriented Offline Reinforcement Learning

Hongchang Zhang, Jianzhun Shao, Yuhang Jiang +2

In offline reinforcement learning, a policy learns to maximize cumulative rewards with a fixed collection of data. Towards conservative strategy, current methods choose to regulari…