2 citations · 2 across the 11 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Yihan Wang, Zhong Guan, Haoran Sun +3
Small language models are attractive backbones for interactive agents, but direct distillation from strong teacher trajectories often turns rich multi-turn behavior into one-shot i…
cs.LG2026
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
Zhong Guan, Yongjian Guo, Haoran Sun +5
Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a c…