4 papers
Finding Kissing Numbers with Game-theoretic Reinforcement Learning
Chengdong Ma, Théo Tao Zhaowei, Pengyu Li +7
Since Isaac Newton first studied the Kissing Number Problem in 1694, determining the maximal number of non-overlapping spheres around a central sphere has remained a defining chall…
Accelerating Robotic Reinforcement Learning with Agent Guidance
Haojun Chen, Zili Zou, Chengdong Ma +4
Reinforcement Learning (RL) offers a powerful paradigm for autonomous robots to master generalist manipulation skills through trial-and-error. However, its real-world application i…
Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
Mingzhi Wang, Chengdong Ma, Qizhi Chen +7
Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF),…
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning
Wenzhe Fan, Zishun Yu, Chengdong Ma +3
In multi-agent reinforcement learning, a commonly considered paradigm is centralized training with decentralized execution. However, in this framework, decentralized execution rest…