works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.LG2026

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Yang Xu, Swetha Ganesh, Vaneet Aggarwal

The paper develops and analyzes model‑free Q‑learning and actor‑critic algorithms that provably learn robust policies for infinite‑horizon average‑reward MDPs under various uncerta…

cs.RO2026

HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation

Junyi Dong, Haotian Luo, Ziwei Xu +11

Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acqui…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

MiniMax, :, Aili Chen +219

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

stat.ML2026

Core-Halo Decomposition: Decentralizing Large-Scale Fixed-Point Problems

Haixiang, Yang Xu, Jiefu Zhang +4

We study solving large-scale fixed-point equation \(x^\star=\bar F(x^\star)\) with decomposition. Standard strict decomposition assigns each agent a disjoint block and evaluates up…

cs.LG2026

Online Bayesian Calibration under Gradual and Abrupt System Changes

Yang Xu, Chiwoo Park

Bayesian model calibration is central to digital twins and computer experiments, as it aligns model outputs with field observations by estimating calibration parameters and correct…

stat.ML2026

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

Yang Xu, Jiefu Zhang, Haixiang Sun +3

Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…