1 citations · 2 across the 22 of their papers we have counts for
29 papers · 1 filter
Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation
Keqin Peng, Chen Li, Yuanxin Ouyang +2
On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxicall…
Beyond Scalar Scores: Exploring LLM-based Metrics for Clinical Significance Evaluation in Radiology Reports
Qingyu Lu, Ruochen Li, Liang Ding +3
Reliable evaluation of generated radiology reports requires strict clinical accuracy, as omitted critical findings or mischaracterized radiographic observations can directly affect…
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…
VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
Jingheng Pan, Xintong Wang, Longyue Wang +3
Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended me…
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
Qihuang Zhong, Liang Ding, Juhua Liu +3
Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data without external supervision. H…
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
Qingyu Lu, Liang Ding, Kanjian Zhang +2
The pursuit of real-time agentic interaction has driven interest in Diffusion-based Large Language Models (dLLMs) as alternatives to auto-regressive backbones, promising to break t…