4 papers · 1 filter
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
Linfeng Cao, Ming Shi, Ness B. Shroff
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and…
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
Ming Shi, Yingbin Liang, Ness B. Shroff +1
Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference la…
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
Amirhossein Roknilamouki, Arnob Ghosh, Ming Shi +3
In Reinforcement Learning (RL), tasks with instantaneous hard constraints present significant challenges, particularly when the decision space is non-convex or non-star-convex. Thi…
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
Linfeng Cao, Ming Shi, Ness B. Shroff
Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences acro…