4 papers
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
Sarat Chandra Bobbili, Ujwal Dinesha, Dheeraj Narasimha +1
Inference-time alignment enables large language models (LLMs) to generate outputs aligned with end-user preferences without further training. Recent post-training methods achieve t…
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
Dheeraj Narasimha, Nicolas Gast
We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real-world systems largely because it resis…
Model Predictive Control is Almost Optimal for Restless Bandit
Nicolas Gast, Dheeraj Narasimha
We consider the discrete time infinite horizon average reward restless markovian bandit (RMAB) problem. We propose a \emph{model predictive control} based non-stationary policy wit…
CONGO: Compressive Online Gradient Optimization
Jeremy Carleton, Prathik Vijaykumar, Divyanshu Saxena +3
We address the challenge of zeroth-order online convex optimization where the objective function's gradient exhibits sparsity, indicating that only a small number of dimensions pos…