4 papers
Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation
Wenhao Zhang
On-policy distillation (OPD) has demonstrated strong empirical gains in enhancing complex reasoning in LLMs by aligning a student model with a teacher's predictive distribution ove…
MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning
Yi Bai, Wenhao Zhang, Yao Chen +3
Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuning data increases, selecting…
Learning Agent-Compatible Context Management for Long-Horizon Tasks
Lu Yi, Runlin Lei, Liuyi Yao +6
LLM agents increasingly face long-horizon tasks such as web search and deep research in real-world applications, where accumulated context can cause long-context degradation and re…
Integrating Chain-of-Thought into Generative Retrieval: A Preliminary Study
Wenhao Zhang, Ruihao Yu, Yi Bai +2
While generative retrieval (GR) demonstrates competitive performance on standard retrieval benchmarks, existing approaches directly map queries to document identifiers (docids) wit…