collaborators

5 papers

cs.LG2026

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

Yongcan Yu, Lingxiao He, Jian Liang +5

Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through…

cs.AI2026

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation

Shuo Lu, Jianjie Cheng, Yinuo Xu +16

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the…

cs.CL2025

DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing

Shuo Lu, Yinuo Xu, Jianjie Cheng +3

Deep Research agents predominantly optimize search policies to maximize retrieval probability. However, we identify a critical bottleneck: the retrieval-utilization gap, where mode…

cs.LG2025

Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning

Yongcan Yu, Lingxiao He, Shuo Lu +10

Recent advances in vision-language models (VLMs) reasoning have been largely attributed to the rise of reinforcement Learning (RL), which has shifted the community's focus away fro…

cs.LG2025

Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances

Shuo Lu, Yingsheng Wang, Lijun Sheng +3

Out-of-distribution (OOD) detection aims to detect test samples outside the training category space, which is an essential component in building reliable machine learning systems.…