2 citations · 2 across the 2 of their papers we have counts for
7 papers
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
Teng Xiao, Yige Yuan, Hamish Ivison +6
This paper introduces MR-Search, an in-context meta reinforcement learning (RL) formulation for agentic search with self-reflection. Instead of optimizing a policy within a single…
Inference-time Alignment in Continuous Space
Yige Yuan, Teng Xiao, Li Yunfan +5
Aligning large language models with human feedback at inference time has received increasing attention due to its flexibility. Existing methods rely on generating multiple response…
Incentivizing Strong Reasoning from Weak Supervision
Yige Yuan, Teng Xiao, Shuchang Tao +4
Large language models (LLMs) have demonstrated impressive performance on reasoning-intensive tasks, but enhancing their reasoning abilities typically relies on either reinforcement…
On a Connection Between Imitation Learning and RLHF
Teng Xiao, Yige Yuan, Mingxiao Li +2
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcem…
Fact-Level Confidence Calibration and Self-Correction
Yige Yuan, Bingbing Xu, Hexiang Tan +5
Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their o…
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
Teng Xiao, Mingxiao Li, Yige Yuan +3
This paper introduces a novel generalized self-imitation learning () framework, which effectively and efficiently aligns large language models with offline demonstra…