8 papers · 1 filter
Selection-Aware Stress Testing for Interactive Agents
Yang Xu, Chenang Li, Jiefu Zhang +3
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We i…
Online Bayesian Calibration under Gradual and Abrupt System Changes
Yang Xu, Chiwoo Park
Bayesian model calibration is central to digital twins and computer experiments, as it aligns model outputs with field observations by estimating calibration parameters and correct…
SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees
Yi Xie, Yangyang Xu, Yi Fan +1
Large language models (LLMs) with a large number of parameters achieve strong performance but are often prohibitively expensive to deploy. Recent work explores using teams of small…
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
Jiefu Zhang, Yang Xu, Vaneet Aggarwal
Navigating safely through dense crowds requires collision avoidance that generalizes beyond the densities seen during training. Learning-based crowd navigation can break under out-…
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
Yang Xu, Swetha Ganesh, Vaneet Aggarwal
We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learni…
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
Yang Xu, Swetha Ganesh, Washim Uddin Mondal +2
This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) with general parametrization. We propose a Primal-Dual Natural Actor-Critic al…