6 papers
Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner
Hao Ma, Zhiqiang Pu, Yang Liu +1
Constraints are essential for stabilizing reinforcement learning fine-tuning (RFT) and preventing degenerate outputs, yet they inherently conflict with the optimization objective b…
Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control
Hao Ma, Zhiqiang Pu, Xiaolin Ai +1
We present GuidedSAC, a novel reinforcement learning (RL) algorithm that facilitates efficient exploration in vast state-action spaces. GuidedSAC leverages large language models (L…
TacEleven: generative tactic discovery for football open play
Siyao Zhao, Hao Ma, Zhiqiang Pu +4
Creating offensive advantages during open play is fundamental to football success. However, due to the highly dynamic and long-sequence nature of open play, the potential tactic sp…
Stochastic Trajectory Prediction under Unstructured Constraints
Hao Ma, Zhiqiang Pu, Shijie Wang +4
Trajectory prediction facilitates effective planning and decision-making, while constrained trajectory prediction integrates regulation into prediction. Recent advances in constrai…
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
Hao Ma, Tianyi Hu, Zhiqiang Pu +4
Reinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominan…
Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
Hao Ma, Shijie Wang, Zhiqiang Pu +2
Guiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward…