works on

From the 2 of 32 linked papers with an AI index.

activity
20242026
collaborators

32 papers

cs.LG2026

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

Naman Saxena, Mudit Gaur, Vaneet Aggarwal

Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hiera…

cs.LG2026

Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

Ankur Naskar, Vaneet Aggarwal

The paper proposes a hierarchical Multilevel Monte Carlo neural critic to reduce bias and cost in actor‑critic reinforcement learning for average‑reward constrained Markov decision…

cs.LG2026

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Yang Xu, Swetha Ganesh, Vaneet Aggarwal

The paper develops and analyzes model‑free Q‑learning and actor‑critic algorithms that provably learn robust policies for infinite‑horizon average‑reward MDPs under various uncerta…

cs.LG2026

Discrete State Diffusion Models: A Sample Complexity Perspective

Aadithya Srikanth, Mudit Gaur, Vaneet Aggarwal

Diffusion models have demonstrated remarkable performance in generating high-dimensional samples across domains such as vision, language, and the sciences. Although continuous-stat…

stat.ML2026

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

Yang Xu, Jiefu Zhang, Haixiang Sun +3

Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…

stat.ML2026

Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients

Yang Xu, Vaneet Aggarwal

We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnosti…