collaborators

5 papers

cs.LG2026

Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries

Linfeng Cao, Ming Shi, Ness B. Shroff

Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and…

cs.LG2026

Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences

Ming Shi, Yingbin Liang, Ness B. Shroff +1

Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference la…

cs.LG2025

Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization

Linfeng Cao, Ming Shi, Ness B. Shroff

Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences acro…

cs.NI2025

Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics

Mohamed A. Abd-Elmagid, Ming Shi, Eylem Ekici +1

We consider a real-time monitoring system where a source node (with energy limitations) aims to keep the information status at a destination node as fresh as possible by scheduling…

cs.LG2025

Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces

Amirhossein Roknilamouki, Arnob Ghosh, Ming Shi +3

In Reinforcement Learning (RL), tasks with instantaneous hard constraints present significant challenges, particularly when the decision space is non-convex or non-star-convex. Thi…