works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.AI2026

Agent-G: Gaussian Guidance for Agentic Reinforcement Learning

Zixuan Wang, Yanrui Miao, Zhengxi Lu +6

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy expl…

cs.CL2026

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Haolei Xu, Xiaowen Xu, Haiwen Hong +5

The paper proposes Relay On-Policy Distillation (Relay-OPD), a method that lets a teacher model temporarily take over a student’s generation when a wrong reasoning prefix is detect…

cs.CL2026

Pause or Fabricate? Training Language Models for Grounded Reasoning

Yiwen Qiu, Linjuan Wu, Yizhou Liu +9

Large language models have achieved remarkable progress on complex reasoning tasks. However, they often implicitly fabricate information when inputs are incomplete, producing confi…

cs.LG2025

Latent Variable Causal Discovery under Selection Bias

Haoyue Dai, Yiwen Qiu, Ignavier Ng +3

Addressing selection bias in latent variable causal discovery is important yet underexplored, largely due to a lack of suitable statistical tools: While various tools beyond basic…

cs.AI2025

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Xingyu Wu, Yuchen Yan, Shangke Lyu +7

Large reasoning models have achieved remarkable performance through extended chain-of-thought sequences, yet this computational freedom leads to excessive token generation even for…

cs.LG2024

Identifying Selections for Unsupervised Subtask Discovery

Yiwen Qiu, Yujia Zheng, Kun Zhang

When solving long-horizon tasks, it is intriguing to decompose the high-level task into subtasks. Decomposing experiences into reusable subtasks can improve data efficiency, accele…