From the 1 of 14 linked papers with an AI index.
14 papers
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
Yang Xu, Swetha Ganesh, Vaneet Aggarwal
The paper develops and analyzes model‑free Q‑learning and actor‑critic algorithms that provably learn robust policies for infinite‑horizon average‑reward MDPs under various uncerta…
HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation
Junyi Dong, Haotian Luo, Ziwei Xu +11
Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acqui…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
Core-Halo Decomposition: Decentralizing Large-Scale Fixed-Point Problems
Haixiang, Yang Xu, Jiefu Zhang +4
We study solving large-scale fixed-point equation \(x^\star=\bar F(x^\star)\) with decomposition. Standard strict decomposition assigns each agent a disjoint block and evaluates up…
Online Bayesian Calibration under Gradual and Abrupt System Changes
Yang Xu, Chiwoo Park
Bayesian model calibration is central to digital twins and computer experiments, as it aligns model outputs with field observations by estimating calibration parameters and correct…
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
Yang Xu, Jiefu Zhang, Haixiang Sun +3
Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…