3 papers
cs.LG2026
Improved Algorithms for Nash Welfare in Linear Bandits
Dhruv Sarkar, Nishant Pandey, Sayak Ray Chowdhury
Nash regret has recently emerged as a principled fairness-aware performance metric for stochastic multi-armed bandits, motivated by the Nash Social Welfare objective. Although this…
cs.LG2025
Revisiting Social Welfare in Bandits: UCB is (Nearly) All You Need
Dhruv Sarkar, Nishant Pandey, Sayak Ray Chowdhury
Regret in stochastic multi-armed bandits traditionally measures the difference between the highest reward and either the arithmetic mean of accumulated rewards or the final reward.…
cs.LG2025
DP-NCB: Privacy Preserving Fair Bandits
Dhruv Sarkar, Nishant Pandey, Sayak Ray Chowdhury
Multi-armed bandit algorithms are fundamental tools for sequential decision-making under uncertainty, with widespread applications across domains such as clinical trials and person…