2 papers
cs.GT2025
Signal Observation Models and Historical Information Integration in Poker Hand Abstraction
Yanchang Fu, Pei Xu, Dongdong Bai +2
Hand abstraction has been instrumental in developing powerful AI for Texas Hold'em poker, a widely studied testbed for imperfect information games (IIGs). Despite its success, the…
cs.LG2024
SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling
Xingzhou Lou, Junge Zhang, Jian Xie +3
Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human pref…