4 citations · 6 across the 9 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
Hanyang Wang, Weijieying Ren, Yuxiang Zhang +4
Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakne…
cs.CL2026
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Yuchen Cai, Ding Cao, Liang Lin +9
On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute this advantage to denser and…