From the 1 of 6 linked papers with an AI index.
6 papers
Bandit PCA with Minimax Optimal Regret
Moïse Blanchard, Dmitrii Ostrovskii, Aadirupa Saha
The paper investigates the bandit-feedback version of online principal component analysis, presenting a new algorithm that achieves near‑optimal regret of order r√(dT) and proving…
One Good Source is All You Need: Near-Optimal Regret for Bandits under Heterogeneous Noise
Amith Bhat, Haipeng Luo, Aadirupa Saha
We study -armed Multiarmed Bandit (MAB) problem with heterogeneous data sources, each exhibiting unknown and distinct noise variances . The learner's obj…
LLM-as-Judge on a Budget
Aadirupa Saha, Aniket Wagde, Branislav Kveton
LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pairs. Since LLM judgments are stoch…
Learning to Allocate Resources with Censored Feedback
Giovanni Montanari, Côme Fiegel, Corentin Pla +2
We study the online resource allocation problem in which at each round, a budget must be allocated across arms under censored feedback. An arm yields a reward if and only i…
Tracking the Best Expert Privately
Aadirupa Saha, Vinod Raman, Hilal Asi
We design differentially private algorithms for the problem of prediction with expert advice under dynamic regret, also known as tracking the best expert. Our work addresses three…
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
Avinandan Bose, Zhihan Xiong, Aadirupa Saha +2
Reinforcement Learning from Human Feedback (RLHF) is currently the leading approach for aligning large language models with human preferences. Typically, these models rely on exten…