works on

From the 2 of 10 linked papers with an AI index.

collaborators

10 papers

cs.GT2026

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Boning Li, Yu Chen, Longbo Huang

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games need…

cs.GT2026

Correlated Chance Sampling for Monte Carlo Counterfactual Regret Minimization

Boning Li, Yu Chen, Longbo Huang

The paper proposes Correlated Chance Sampling (CCS-MCCFR), a modification to Monte Carlo Counterfactual Regret Minimization that uses persistent randomized streams to reduce sampli…

cs.LG2026

On the Sublinear Regret of Continuous K-Max Bandits

Yu Chen, Siwei Wang, Longbo Huang +1

The paper studies continuous K‑max combinatorial multi‑armed bandits, proposing the DCK‑UCB algorithm with adaptive discretization and bias‑corrected confidence bounds that achieve…

cs.LG2026

When Context Returns: Toward Robust Internalization in On-Policy Distillation

Xun Wang, Ruishuo Chen, Zhuoran Li +2

Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…

cs.CL2026

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

Ruishuo Chen, Yu Chen, Zhuoran Li +1

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without…

cs.LG2026

Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes

Yu Chen, Yuhao Liu, Jiatai Huang +2

We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity…