Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
Shulun Chen, Runlong Zhou, Zihan Zhang +2
Aligning large language models (LLMs) with human preferences has proven effective for enhancing model capabilities, yet standard preference modeling using the Bradley-Terry model a…
cs.LG2025
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
Shulun Chen, Runlong Zhou, Zihan Zhang +2
We consider the gap-dependent regret bounds for episodic MDPs. We show that the Monotonic Value Propagation (MVP) algorithm achieves a variance-aware gap-dependent regret bound of…