papers

Publications (21)

cs.LG2021

L2E: Learning to Exploit Your Opponent

Zhe Wu, Kai Li, Enmin Zhao +5

Opponent modeling is essential to exploit sub-optimal opponents in strategic interactions. Most previous works focus on building explicit models to directly predict the opponents'…

cs.GT2023

Policy Space Diversity for Non-Transitive Games

Jian Yao, Weiming Liu, Haobo Fu +4

Policy-Space Response Oracles (PSRO) is an influential algorithm framework for approximating a Nash Equilibrium (NE) in multi-agent non-transitive games. Many previous studies have…

cs.MA2025

Maximum Entropy Heterogeneous-Agent Reinforcement Learning

Jiarong Liu, Yifan Zhong, Siyi Hu +4

Multi-agent reinforcement learning (MARL) has been shown effective for cooperative games in recent years. However, existing state-of-the-art methods face challenges related to samp…

cs.LG2024

Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning

Hanlin Yang, Jian Yao, Weiming Liu +13

Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory,…

cs.LG2025

Deep (Predictive) Discounted Counterfactual Regret Minimization

Hang Xu, Kai Li, Haobo Fu +3

Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers u…

cs.LG2024

Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent

Hang Xu, Kai Li, Bingyun Liu +4

Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. It decomposes the total regret into counterfactual regrets,…

cs.LG2018

Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Jiechao Xiong, Qing Wang, Zhuoran Yang +7

Most existing deep reinforcement learning (DRL) frameworks consider either discrete action space or continuous action space solely. Motivated by applications in computer games, we…

cs.LG2025

Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning

Jinmin He, Kai Li, Yifan Zang +4

Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any onlin…

cs.AI2025

TacticCraft: Natural Language-Driven Tactical Adaptation for StarCraft II

Weiyu Ma, Jiwen Jiang, Haobo Fu +1

We present an adapter-based approach for tactical conditioning of StarCraft II AI agents. Current agents, while powerful, lack the ability to adapt their strategies based on high-l…

cs.AI2024

Reaching Consensus in Cooperative Multi-Agent Reinforcement Learning with Goal Imagination

Liangzhou Wang, Kaiwen Zhu, Fengming Zhu +6

Reaching consensus is key to multi-agent coordination. To accomplish a cooperative task, agents need to coherently select optimal joint actions to maximize the team reward. However…

cs.LG2026

AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization

Shengda Gu, Kai Li, Xinyi Ke +3

AutoPref uses a large language model to automatically discover and compose pairwise loss and weighting programs that define preference objectives for neural combinatorial optimizat…

#neural combinatorial optimization#preference learning#reinforcement learning#automated objective discovery
cs.AI2024

Enhance Reasoning for Large Language Models in the Game Werewolf

Shuang Wu, Liwen Zhu, Tao Yang +4

This paper presents an innovative framework that integrates Large Language Models (LLMs) with an external Thinker module to enhance the reasoning capabilities of LLM-based agents.…

cs.GT2024

Enhanced Equilibria-Solving via Private Information Pre-Branch Structure in Adversarial Team Games

Chen Qiu, Haobo Fu, Kai Li +3

In ex ante coordinated adversarial team games (ATGs), a team competes against an adversary, and the team members are only allowed to coordinate their strategies before the game sta…

cs.AI2025

AI Deception: Risks, Dynamics, and Controls

Boyuan Chen, Sitong Fang, Jiaming Ji +56

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an…

cs.LG2025

Diversity from Human Feedback

Ren-Jian Wang, Ke Xue, Yutong Wang +4

Diversity plays a significant role in many problems, such as ensemble learning, reinforcement learning, and combinatorial optimization. How to define the diversity measure is a lon…

cs.LG2025

Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance

Jinmin He, Kai Li, Yifan Zang +4

Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing appr…

cs.NE2024

Heterogeneous Multi-agent Zero-Shot Coordination by Coevolution

Ke Xue, Yutong Wang, Cong Guan +5

Generating agents that can achieve zero-shot coordination (ZSC) with unseen partners is a new challenge in cooperative multi-agent reinforcement learning (MARL). Recently, some stu…

cs.NE2024

Pointer Networks Trained Better via Evolutionary Algorithms

Muyao Zhong, Shengcai Liu, Bingdong Li +3

Pointer Network (PtrNet) is a specific neural network for solving Combinatorial Optimization Problems (COPs). While PtrNets offer real-time feed-forward inference for complex COPs…

cs.AI2026

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies

Guangyu Zhao, Kewei Lian, Haoxuan Ru +8

Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the cho…

cs.LG2025

SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization

Zhengcheng Wang, Zichuan Lin, Yijun Yang +2

Existing Vision-Language Navigation (VLN) agents based on Large Vision-Language Models (LVLMs) often suffer from perception errors, reasoning errors, and planning errors, which sig…

cs.AI2024

Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing

Jinmin He, Kai Li, Yifan Zang +4

Multi-task reinforcement learning endeavors to accomplish a set of different tasks with a single policy. To enhance data efficiency by sharing parameters across multiple tasks, a c…