9 papers
On the Geometry of On-Policy Distillation
Zhennan Shen, Yanshu Li, Qingyu Yin +6
On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of O…
A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction
Chengguang Gan, Sunbowen Lee, Qingyu Yin +9
The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improve each other when jointly modele…
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
Chak Tou Leong, Dingwei Chen, Heming Xia +4
Large reasoning models (LRMs) have achieved remarkable success in complex problem-solving, yet they often suffer from computational redundancy or reasoning unfaithfulness. Current…
Evaluating Parameter Efficient Methods for RLVR
Qingyu Yin, Yulun Wu, Zhennan Shen +6
We systematically evaluate Parameter-Efficient Fine-Tuning (PEFT) methods under the paradigm of Reinforcement Learning with Verifiable Rewards (RLVR). RLVR incentivizes language mo…
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
Shuxun Wang, Qingyu Yin, Chak Tou Leong +2
Repetition curse is a phenomenon where Large Language Models (LLMs) generate repetitive sequences of tokens or cyclic sequences. While the repetition curse has been widely observed…
Probing the Difficulty Perception Mechanism of Large Language Models
Sunbowen Lee, Qingyu Yin, Chak Tou Leong +5
Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an es…