activity
20242026
collaborators

6 papers

cs.LG2026

GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning

Zhiyuan Fan, Gabriele Farina

Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial opponents, necessitating st…

cs.LG2026

Online Learning and Equilibrium Computation with Ranking Feedback

Mingyang Liu, Yongshan Chen, Zhiyuan Fan +3

Online learning in arbitrary, and possibly adversarial, environments has been extensively studied in sequential decision-making, and it is closely connected to equilibrium computat…

cs.LG2026

Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning

Zichao Li, Jie Lou, Fangchen Dong +8

Reinforcement learning significantly enhances LLM capabilities but suffers from a critical issue: length inflation, where models adopt verbosity or inefficient reasoning to maximiz…

cs.LG2025

On the Universal Near Optimality of Hedge in Combinatorial Settings

Zhiyuan Fan, Arnab Maiti, Kevin Jamieson +2

In this paper, we study the classical Hedge algorithm in combinatorial settings. In each round, the learner selects a vector from a set ,…

cs.LG2025

Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries

Arnab Maiti, Zhiyuan Fan, Kevin Jamieson +2

In this paper, we study the online shortest path problem in directed acyclic graphs (DAGs) under bandit feedback against an adaptive adversary. Given a DAG with a sour…

cs.LG2024

On the Optimality of Dilated Entropy and Lower Bounds for Online Learning in Extensive-Form Games

Zhiyuan Fan, Christian Kroer, Gabriele Farina

First-order methods (FOMs) are arguably the most scalable algorithms for equilibrium computation in large extensive-form games. To operationalize these methods, a distance-generati…