3 papers
cs.AI2025
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
Sarat Chandra Bobbili, Ujwal Dinesha, Dheeraj Narasimha +1
Inference-time alignment enables large language models (LLMs) to generate outputs aligned with end-user preferences without further training. Recent post-training methods achieve t…
cs.LG2025
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
Guojun Xiong, Ujwal Dinesha, Debajoy Mukherjee +2
Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Marko…
cs.AI2025
Risk-Averse Finetuning of Large Language Models
Sapana Chaudhary, Ujwal Dinesha, Dileep Kalathil +1
We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-…