3 papers
cs.LG2026
VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
Pengcheng Li, Zhengyang Zhang, Dongxu Zhang +2
Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable termin…
cs.AI2026
CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
De Jiang, Zhengyang Zhang, Kehong Yuan +1
Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy sourc…
cs.LG2026
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
De Jiang, Zhengyang Zhang, Kehong Yuan +1
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this…