25 papers
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
Linfeng Cao, Ming Shi, Ness B. Shroff
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and…
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
Yuchen Liang, Ness Shroff, Yingbin Liang
Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, but, especially for uniform-rate models, they often require many steps to g…
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
Ming Shi, Yingbin Liang, Ness B. Shroff +1
Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference la…
An LP-based Sampling Policy for Multi-Armed Bandits with Side-Observations and Stochastic Availability
Ashutosh Soni, Peizhong Ju, Atilla Eryilmaz +1
We study the stochastic multi-armed bandit (MAB) problem where an underlying network structure enables side-observations across related actions. We use a bipartite graph to link ac…
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
Amirhossein Roknilamouki, Arnob Ghosh, Eylem Ekici +1
While offline reinforcement learning provides reliable policies for real-world deployment, its inherent pessimism severely restricts an agent's ability to explore and collect novel…
Sharp Convergence Rates for Masked Diffusion Models
Yuchen Liang, Zhiheng Tan, Ness Shroff +1
Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, with masked (absorbing-rate) variants emerging as competitive alternatives…