activity
20242026
collaborators

5 papers

cs.LG2026

Finding Kissing Numbers with Game-theoretic Reinforcement Learning

Chengdong Ma, Théo Tao Zhaowei, Pengyu Li +7

Since Isaac Newton first studied the Kissing Number Problem in 1694, determining the maximal number of non-overlapping spheres around a central sphere has remained a defining chall…

cs.RO2026

Accelerating Robotic Reinforcement Learning with Agent Guidance

Haojun Chen, Zili Zou, Chengdong Ma +4

Reinforcement Learning (RL) offers a powerful paradigm for autonomous robots to master generalist manipulation skills through trial-and-error. However, its real-world application i…

cs.CL2025

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Mingzhi Wang, Chengdong Ma, Qizhi Chen +7

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF),…

cs.MA2024

Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning

Wenzhe Fan, Zishun Yu, Chengdong Ma +3

In multi-agent reinforcement learning, a commonly considered paradigm is centralized training with decentralized execution. However, in this framework, decentralized execution rest…

cs.GT2024

Sample-Efficient Regret-Minimizing Double Oracle in Extensive-Form Games

Xiaohang Tang, Chiyuan Wang, Chengdong Ma +3

Extensive-Form Game (EFG) represents a fundamental model for analyzing sequential interactions among multiple agents and the primary challenge to solve it lies in mitigating sample…