activity
20242026
collaborators

7 papers

cs.GT2026

Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles

Jiesong Lian, Yucong Huang, Chengdong Ma +4

For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown th…

cs.AI2025

Roadmap on Incentive Compatibility for AI Alignment and Governance in Sociotechnical Systems

Zhaowei Zhang, Fengshuo Bai, Mingzhi Wang +3

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides h…

cs.CL2025

Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

Zhaowei Zhang, Fengshuo Bai, Qizhi Chen +5

How to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, chang…

cs.RO2025

Falcon: Fast Visuomotor Policies via Partial Denoising

Haojun Chen, Minghao Liu, Chengdong Ma +8

Diffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for ac…

cs.CL2025

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Mingzhi Wang, Chengdong Ma, Qizhi Chen +7

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF),…

cs.GT2024

Conflux-PSRO: Effectively Leveraging Collective Advantages in Policy Space Response Oracles

Yucong Huang, Jiesong Lian, Mingzhi Wang +2

Policy Space Response Oracle (PSRO) with policy population construction has been demonstrated as an effective method for approximating Nash Equilibrium (NE) in zero-sum games. Exis…