5 papers
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
Linfeng Cao, Ming Shi, Ness B. Shroff
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and…
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
Ming Shi, Yingbin Liang, Ness B. Shroff +1
Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference la…
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
Linfeng Cao, Ming Shi, Ness B. Shroff
Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences acro…
Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics
Mohamed A. Abd-Elmagid, Ming Shi, Eylem Ekici +1
We consider a real-time monitoring system where a source node (with energy limitations) aims to keep the information status at a destination node as fresh as possible by scheduling…
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
Amirhossein Roknilamouki, Arnob Ghosh, Ming Shi +3
In Reinforcement Learning (RL), tasks with instantaneous hard constraints present significant challenges, particularly when the decision space is non-convex or non-star-convex. Thi…