activity
20242026
most citedFact-Level Confidence Calibration and Self-Correction

2 citations · 2 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Meta-Reinforcement Learning with Self-Reflection for Agentic Search

Teng Xiao, Yige Yuan, Hamish Ivison +6

This paper introduces MR-Search, an in-context meta reinforcement learning (RL) formulation for agentic search with self-reflection. Instead of optimizing a policy within a single…

cs.CL2025

Inference-time Alignment in Continuous Space

Yige Yuan, Teng Xiao, Li Yunfan +5

Aligning large language models with human feedback at inference time has received increasing attention due to its flexibility. Existing methods rely on generating multiple response…

cs.CL2025

Incentivizing Strong Reasoning from Weak Supervision

Yige Yuan, Teng Xiao, Shuchang Tao +4

Large language models (LLMs) have demonstrated impressive performance on reasoning-intensive tasks, but enhancing their reasoning abilities typically relies on either reinforcement…

cs.LG2025

On a Connection Between Imitation Learning and RLHF

Teng Xiao, Yige Yuan, Mingxiao Li +2

This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcem…

cs.CL20242 cited

Fact-Level Confidence Calibration and Self-Correction

Yige Yuan, Bingbing Xu, Hexiang Tan +5

Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their o…

cs.CL2024

How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

Teng Xiao, Mingxiao Li, Yige Yuan +3

This paper introduces a novel generalized self-imitation learning () framework, which effectively and efficiently aligns large language models with offline demonstra…