activity
20182025
most citedOn Solving Minimax Optimization Locally: A Follow-the-Ridge Approach

18 citations · 33 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG2025★ 1 cited

Is Elo Rating Reliable? A Study Under Model Misspecification

Shange Tang, Yuanhao Wang, Chi Jin

Elo rating, widely used for skill assessment across diverse domains ranging from competitive games to large language models, is often understood as an incremental update algorithm…

cs.LG2023★ 1 cited

Is RLHF More Difficult than Standard RL?

Yuanhao Wang, Qinghua Liu, Chi Jin

Reinforcement learning from Human Feedback (RLHF) learns from preference signals, while standard Reinforcement Learning (RL) directly learns from reward signals. Preferences arguab…

cs.LG2023★ 2 cited

Breaking the Curse of Multiagency: Provably Efficient Decentralized Multi-Agent RL with Function Approximation

Yuanhao Wang, Qinghua Liu, Yu Bai +1

A unique challenge in Multi-Agent Reinforcement Learning (MARL) is the curse of multiagency, where the description length of the game as well as the complexity of many existing lea…

cs.LG2022

Learning Rationalizable Equilibria in Multiplayer Games

Yuanhao Wang, Dingwen Kong, Yu Bai +1

A natural goal in multiagent learning besides finding equilibria is to learn rationalizable behavior, where players learn to avoid iteratively dominated actions. However, even in t…

cs.LG2022

Learning Markov Games with Adversarial Opponents: Efficient Algorithms and Fundamental Limits

Qinghua Liu, Yuanhao Wang, Chi Jin

An ideal strategy in zero-sum games should not only grant the player an average reward no less than the value of Nash equilibrium, but also exploit the (adaptive) opponents when th…

cs.LG2021★ 11 cited

V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL

Chi Jin, Qinghua Liu, Yuanhao Wang +1

A major challenge of multiagent reinforcement learning (MARL) is the curse of multiagents, where the size of the joint action space scales exponentially with the number of agents.…