2 papers
cs.LG2026
MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection
Jeremy Carleton, Debajoy Mukherjee, Srinivas Shakkottai +1
Large Language Models (LLMs) are increasingly deployed across diverse applications that demand balancing multiple, often conflicting, objectives -- such as helpfulness, harmlessnes…
cs.LG2025
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
Guojun Xiong, Ujwal Dinesha, Debajoy Mukherjee +2
Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Marko…