works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.LG2026

Bandit PCA with Minimax Optimal Regret

Moïse Blanchard, Dmitrii Ostrovskii, Aadirupa Saha

The paper investigates the bandit-feedback version of online principal component analysis, presenting a new algorithm that achieves near‑optimal regret of order r√(dT) and proving…

cs.LG2026

One Good Source is All You Need: Near-Optimal Regret for Bandits under Heterogeneous Noise

Amith Bhat, Haipeng Luo, Aadirupa Saha

We study -armed Multiarmed Bandit (MAB) problem with heterogeneous data sources, each exhibiting unknown and distinct noise variances . The learner's obj…

cs.LG2026

LLM-as-Judge on a Budget

Aadirupa Saha, Aniket Wagde, Branislav Kveton

LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pairs. Since LLM judgments are stoch…

cs.LG2026

Learning to Allocate Resources with Censored Feedback

Giovanni Montanari, Côme Fiegel, Corentin Pla +2

We study the online resource allocation problem in which at each round, a budget must be allocated across arms under censored feedback. An arm yields a reward if and only i…

cs.LG2025

Tracking the Best Expert Privately

Aadirupa Saha, Vinod Raman, Hilal Asi

We design differentially private algorithms for the problem of prediction with expert advice under dynamic regret, also known as tracking the best expert. Our work addresses three…

cs.LG2024

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

Avinandan Bose, Zhihan Xiong, Aadirupa Saha +2

Reinforcement Learning from Human Feedback (RLHF) is currently the leading approach for aligning large language models with human preferences. Typically, these models rely on exten…