5 papers
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
Chen He, Yuhao Wu, Lei Wang +2
Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcome…
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
Xun Jiang, Yufan Gu, Disen Hu +5
Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues…
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
Shenshen Li, Xing Xu, Kaiyuan Deng +3
While multi-modal large language models (MLLMs) have made significant progress in complex reasoning tasks via reinforcement learning, it is commonly believed that extensive trainin…
What Makes Reasoning Invalid: Echo Reflection Mitigation for Large Language Models
Chen He, Xun Jiang, Lei Wang +5
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of reasoning tasks. Recent methods have further improved LLM performance in complex mathem…
Conformal Lesion Segmentation for 3D Medical Images
Binyu Tan, Zhiyuan Wang, Jinhao Duan +4
Medical image segmentation serves as a critical component of precision medicine, enabling accurate localization and delineation of pathological regions, such as lesions. However, e…