3 papers
cs.LG2025
ELO-Rated Sequence Rewards: Advancing Reinforcement Learning Models
Qi Ju, Falin Hei, Zhemei Fang +1
Reinforcement Learning (RL) heavily relies on the careful design of the reward function. However, accurately assigning rewards to each state-action pair in Long-Term Reinforcement…
cs.GT2025
Preference-CFR Beyond Nash Equilibrium for Better Game Strategies
Qi Ju, Thomas Tellier, Meng Sun +2
Artificial intelligence (AI) has surpassed top human players in a variety of games. In imperfect information games, these achievements have primarily been driven by Counterfactual…
cs.AI2024
Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play
Ju Qi, Falin Hei, Ting Feng +3
Counterfactual Regret Minimization (CFR) and its variants are widely recognized as effective algorithms for solving extensive-form imperfect information games. Recently, many impro…