Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A Formula-Driven Survey and Research Agenda for On-Policy Distillation
Bowen Zhang
On-policy distillation (OPD) trains an LLM on states induced by the current or recent student policy: the student generates complete or partial rollouts, a teacher or self-teacher…
cs.AI2026
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
Jiazheng Kang, Bowen Zhang, Zixin Song +4
ReAct-style agents for search-intensive, multi-step reasoning tasks rely largely on their own internal judgment to decide what evidence to seek, which reasoning or action step to t…