From the 2 of 10 linked papers with an AI index.
10 papers
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
Boning Li, Yu Chen, Longbo Huang
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games need…
Correlated Chance Sampling for Monte Carlo Counterfactual Regret Minimization
Boning Li, Yu Chen, Longbo Huang
The paper proposes Correlated Chance Sampling (CCS-MCCFR), a modification to Monte Carlo Counterfactual Regret Minimization that uses persistent randomized streams to reduce sampli…
On the Sublinear Regret of Continuous K-Max Bandits
Yu Chen, Siwei Wang, Longbo Huang +1
The paper studies continuous K‑max combinatorial multi‑armed bandits, proposing the DCK‑UCB algorithm with adaptive discretization and bias‑corrected confidence bounds that achieve…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Xun Wang, Ruishuo Chen, Zhuoran Li +2
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Ruishuo Chen, Yu Chen, Zhuoran Li +1
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without…
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
Yu Chen, Yuhao Liu, Jiatai Huang +2
We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity…