works on

From the 1 of 15 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Yang Xu, Swetha Ganesh, Vaneet Aggarwal

The paper develops and analyzes model‑free Q‑learning and actor‑critic algorithms that provably learn robust policies for infinite‑horizon average‑reward MDPs under various uncerta…

cs.LG2026

Online Bayesian Calibration under Gradual and Abrupt System Changes

Yang Xu, Chiwoo Park

Bayesian model calibration is central to digital twins and computer experiments, as it aligns model outputs with field observations by estimating calibration parameters and correct…

cs.LG2026

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees

Yi Xie, Yangyang Xu, Yi Fan +1

Large language models (LLMs) with a large number of parameters achieve strong performance but are often prohibitively expensive to deploy. Recent work explores using teams of small…

cs.LG2026

Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds

Jiefu Zhang, Yang Xu, Vaneet Aggarwal

Navigating safely through dense crowds requires collision avoidance that generalizes beyond the densities seen during training. Learning-based crowd navigation can break under out-…

cs.LG2025

Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm

Yang Xu, Swetha Ganesh, Washim Uddin Mondal +2

This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) with general parametrization. We propose a Primal-Dual Natural Actor-Critic al…

cs.LG2025

Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes

Bhargav Ganguly, Yang Xu, Vaneet Aggarwal

This paper investigates the potential of quantum acceleration in addressing infinite horizon Markov Decision Processes (MDPs) to enhance average reward outcomes. We introduce an in…