3 citations · 4 across the 2 of their papers we have counts for
5 papers
Preference-based Reinforcement Learning with Finite-Time Guarantees
Yichong Xu, Ruosong Wang, Lin F. Yang +2
Preference-based Reinforcement Learning (PbRL) replaces reward values in traditional reinforcement learning by preferences to better elicit human opinion on the target objective, e…
Zeroth Order Non-convex optimization with Dueling-Choice Bandits
Yichong Xu, Aparna Joshi, Aarti Singh +1
We consider a novel setting of zeroth order non-convex optimization, where in addition to querying the function value at a given point, we can also duel two points and get the poin…
Thresholding Bandit Problem with Both Duels and Pulls
Yichong Xu, Xi Chen, Aarti Singh +1
The Thresholding Bandit Problem (TBP) aims to find the set of arms with mean rewards greater than a given threshold. We consider a new setting of TBP, where in addition to pulling…
Regression with Comparisons: Escaping the Curse of Dimensionality with Ordinal Information
Yichong Xu, Sivaraman Balakrishnan, Aarti Singh +1
In supervised learning, we typically leverage a fully labeled dataset to design methods for function estimation or prediction. In many practical situations, we are able to obtain a…
Noise-Tolerant Interactive Learning from Pairwise Comparisons
Yichong Xu, Hongyang Zhang, Aarti Singh +2
We study the problem of interactively learning a binary classifier using noisy labeling and pairwise comparison oracles, where the comparison oracle answers which one in the given…