4 papers
Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
Chia-Hsuan Lee, Mingyang Zhou, Renkun Ni +6
Preference optimization methods such as DPO and KTO are widely used for aligning language models, yet little is understood about what properties of preference data drive downstream…
Benchmarking drug-drug interaction prediction methods: a perspective of distribution changes
Zhenqian Shen, Mingyang Zhou, Yongqi Zhang +1
Motivation: Emerging drug-drug interaction (DDI) prediction is crucial for new drugs but is hindered by distribution changes between known and new drugs in real-world scenarios. Cu…
Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
Mingyang Zhou, Quanming Yao, Lun Du +2
Reproducing machine learning papers is essential for scientific progress but remains challenging for both humans and automated agents. Existing agent-based methods often struggle t…