activity
20242026
collaborators

7 papers

cs.CL2026

Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation

Kaiyuan Liu, Ziyuan Zhuang, Yang Bai +3

On-policy distillation (OPD) trains a student model on its own rollouts using dense feedback from a stronger teacher. Prior literature suggests that, provided teacher feedback is a…

cs.LG2026

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization

Yang Bai, Kaiyuan Liu, Ziyuan Zhuang +5

Complex reinforcement learning environments frequently employ multi-task and mixed-reward formulations. In these settings, heterogeneous reward distributions and correlated reward…

cs.CV2026

Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks

Jing Jin, Hao Liu, Yan Bai +9

Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains challenging. STEM reasoning is a…

cs.CL2025

A Survey on LLM Mid-Training

Chengying Tu, Xuemiao Zhang, Rongxiang Weng +6

Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage…

cs.AI2025

AdaR: A Framework for Equipping LLMs with Adaptive Reasoning

Zhejian Lai, Xiang Geng, Zhijun Wang +7

Mathematical reasoning is a primary indicator of large language models (LLMs) intelligence. However, existing LLMs exhibit failures in robustness and generalization. This paper att…

cs.CL2025

Libra: Assessing and Improving Reward Model by Learning to Think

Meng Zhou, Bei Li, Jiahao Liu +5

Reinforcement learning (RL) has significantly improved the reasoning ability of large language models. However, current reward models underperform in challenging reasoning scenario…