activity
20242026
collaborators

6 papers

cs.LG2026

DART: aDaptive Accept RejecT for non-linear top-K subset identification

Mridul Agarwal, Vaneet Aggarwal, Christopher J. Quinn +1

We consider the bandit problem of selecting out of arms at each time step. The reward can be a non-linear function of the rewards of the selected individual arms. The direc…

cs.LG2025

Hierarchical Deep Counterfactual Regret Minimization

Jiayu Chen, Zhekai Wang, Vaneet Aggarwal

Imperfect Information Games (IIGs) offer robust models for scenarios where decision-makers face uncertainty or lack complete information. Counterfactual Regret Minimization (CFR) h…

cs.LG2025

Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm

Qinbo Bai, Mridul Agarwal, Vaneet Aggarwal

Many engineering problems have multiple objectives, and the overall aim is to optimize a non-linear function of these objectives. In this paper, we formulate the problem of maximiz…

cs.LG2025

Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback

Mohammad Pedramfar, Vaneet Aggarwal

This paper investigates the problem of combinatorial multiarmed bandits with stochastic submodular (in expectation) rewards and full-bandit delayed feedback, where the delayed feed…

cs.LG2025

Q-GADMM: Quantized Group ADMM for Communication Efficient Decentralized Machine Learning

Anis Elgabli, Jihong Park, Amrit S. Bedi +3

In this article, we propose a communication-efficient decentralized machine learning (ML) algorithm, coined quantized group ADMM (Q-GADMM). To reduce the number of communication li…

cs.NI2024

Learning-based Two-tiered Online Optimization of Region-wide Datacenter Resource Allocation

Chang-Lin Chen, Hanhan Zhou, Jiayu Chen +8

Online optimization of resource management for large-scale data centers and infrastructures to meet dynamic capacity reservation demands and various practical constraints (e.g., fe…