activity
20242026
most citedSelf-Clustering Hierarchical Multi-Agent Reinforcement Learning with Extensible Cooperation Graph

1 citations · 3 across the 10 of their papers we have counts for

collaborators

10 papers

cs.LG2026

Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner

Hao Ma, Zhiqiang Pu, Yang Liu +1

Constraints are essential for stabilizing reinforcement learning fine-tuning (RFT) and preventing degenerate outputs, yet they inherently conflict with the optimization objective b…

cs.LG2026

Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control

Hao Ma, Zhiqiang Pu, Xiaolin Ai +1

We present GuidedSAC, a novel reinforcement learning (RL) algorithm that facilitates efficient exploration in vast state-action spaces. GuidedSAC leverages large language models (L…

cs.MA2025

Heterogeneity in Multi-Agent Reinforcement Learning

Tianyi Hu, Zhiqiang Pu, Yuan Wang +3

Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy…

cs.LG2025

CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning

Jinyuan Feng, Chaopeng Wei, Tenghai Qiu +2

In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has bee…

cs.AI20251 cited

Unreal-MAP: Unreal-Engine-Based General Platform for Multi-Agent Reinforcement Learning

Tianyi Hu, Qingxu Fu, Zhiqiang Pu +2

In this paper, we propose Unreal Multi-Agent Playground (Unreal-MAP), an MARL general platform based on the Unreal-Engine (UE). Unreal-MAP allows users to freely create multi-agent…

cs.RO2025

Stochastic Trajectory Prediction under Unstructured Constraints

Hao Ma, Zhiqiang Pu, Shijie Wang +4

Trajectory prediction facilitates effective planning and decision-making, while constrained trajectory prediction integrates regulation into prediction. Recent advances in constrai…