2 papers
cs.LG2026
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a…
cs.AI2026
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
Ashutosh Hathidara, Julien Yu, Vaishali Senthil +2
Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-use…