collaborators

11 papers

cs.AI2026

Retrieval Grounding Latent Reasoning for Dense Retrieval

Gang Zhou, Xiongxi Yu, Hu Tian +5

Reasoning-intensive retrieval requires text representations to capture not only semantic similarity, but also the reasoning needed to determine relevance under a given retrieval in…

cs.AI2026

Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation

Shuaitao Zhao, Feng Ni, Lichao Ma +4

The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new challenges for compliance detec…

cs.AI2026

How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning

Lichao Ma, Yang Sun, Shuaitao Zhao +9

Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its t…

cs.LG2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Yi Yang, Cong Qin, Xiaodan Liu +8

Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens…

cs.CL2026

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

Chishui Chen, Yaoyou Fan, Te Sun +11

On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn age…

cs.CL2026

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

Jingwen Chen, Wenkai Yang, Shengda Fan +7

Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward continual learning in large l…