3 papers
cs.AI2026
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Haolong Qian, Xianliang Yang, Yinuo ma +6
Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher rewar…
cs.CL2026
Unsupervised Text Style Transfer for Controllable Intensity
Shuhuan Gu, Wenbiao Tao, Xinchen Ma +4
Unsupervised Text Style Transfer (UTST) aims to build a system to transfer the stylistic properties of a given text without parallel text pairs. Compared with text transfer between…
cs.CV2025
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…