2 papers
cs.LG2026
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
Shengbo Wang
We study reinforcement learning in infinite-horizon discounted Markov decision processes with continuous state spaces, where data are generated online from a single trajectory unde…
cs.LG2025
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
Shengbo Wang, Hong Sun, Ke Li
Interactive preference elicitation (IPE) aims to substantially reduce human effort while acquiring human preferences in wide personalization systems. Dueling bandit (DB) algorithms…