collaborators

7 papers

cs.LG2026

A Simple Reduction Scheme for Constrained Contextual Bandits with Adversarial Contexts via Regression

Dhruv Sarkar, Abhishek Sinha

We study constrained contextual bandits (CCB) with adversarially chosen contexts, where each action yields a random reward and incurs a random cost. We adopt the standard realizabi…

cs.LG2025

Optimal Anytime Algorithms for Online Convex Optimization with Adversarial Constraints

Dhruv Sarkar, Abhishek Sinha

We propose an anytime online algorithm for the problem of learning a sequence of adversarial convex cost functions while approximately satisfying another sequence of adversarial on…

cs.LG2025

Revisiting Social Welfare in Bandits: UCB is (Nearly) All You Need

Dhruv Sarkar, Nishant Pandey, Sayak Ray Chowdhury

Regret in stochastic multi-armed bandits traditionally measures the difference between the highest reward and either the arithmetic mean of accumulated rewards or the final reward.…

cs.LG2025

Online Learning for Approximately-Convex Functions with Long-term Adversarial Constraints

Dhruv Sarkar, Samrat Mukhopadhyay, Abhishek Sinha

We study an online learning problem with long-term budget constraints in the adversarial setting. In this problem, at each round , the learner selects an action from a convex de…

cs.LG2025

DP-NCB: Privacy Preserving Fair Bandits

Dhruv Sarkar, Nishant Pandey, Sayak Ray Chowdhury

Multi-armed bandit algorithms are fundamental tools for sequential decision-making under uncertainty, with widespread applications across domains such as clinical trials and person…

cs.CV2025

TAPS : Frustratingly Simple Test Time Active Learning for VLMs

Dhruv Sarkar, Aprameyo Chakrabartty, Bibhudatta Bhanja

Test-Time Optimization enables models to adapt to new data during inference by updating parameters on-the-fly. Recent advances in Vision-Language Models (VLMs) have explored learni…