works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.AI2026

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Zhengyu Chen, Teng Xiao, Huaisheng Zhu +3

Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, eval…

cs.AI2026

Rethinking the Evaluation of Harness Evolution for Agents

Yike Wang, Huaisheng Zhu, Zhengyu Hu +7

The paper reexamines how automatic harness evolution for large language model agents is evaluated, comparing it to simple test‑time scaling baselines and finding that it offers lim…

cs.LG2026

Meta-Reinforcement Learning with Self-Reflection for Agentic Search

Teng Xiao, Yige Yuan, Hamish Ivison +6

This paper introduces MR-Search, an in-context meta reinforcement learning (RL) formulation for agentic search with self-reflection. Instead of optimizing a policy within a single…

cs.CL2026

Inference-time Alignment in Continuous Space

Yige Yuan, Teng Xiao, Li Yunfan +5

Aligning large language models with human feedback at inference time has received increasing attention due to its flexibility. Existing methods rely on generating multiple response…

cs.CL2026

Incentivizing Strong Reasoning from Weak Supervision

Yige Yuan, Teng Xiao, Shuchang Tao +4

Large language models (LLMs) have demonstrated impressive performance on reasoning-intensive tasks, but enhancing their reasoning abilities typically relies on either reinforcement…

cs.LG2025

On a Connection Between Imitation Learning and RLHF

Teng Xiao, Yige Yuan, Mingxiao Li +2

This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcem…