3 papers
cs.CL2026
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Rongcan Pei, Zhepei Wei, Shuyao Xu +3
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging pri…
cs.IR2026
Sci-Surf: Navigating Scientific Literature Discovery through Human Feedback and Intelligent Summarization
Fang Guo, Qi Zhu, Rongcan Pei +3
The rapid growth of scientific publications makes it increasingly difficult for researchers to identify relevant new studies and effectively comprehend them. Existing academic disc…
cs.CV2026
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
Rongcan Pei, Huan Li, Fang Guo +1
While Vision-Language Models (VLMs) have shown promise in textual understanding, they face significant challenges when handling long context and complex reasoning tasks. In this pa…