collaborators

6 papers

cs.AI2026

The Illusion of : Evaluating the Breakdown of Counterfactual Reasoning in LLMs

Yucheng Wang, Yuetian Du, Zhengyi Liu +8

Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks large…

cs.CV2026

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

Yuetian Du, Yucheng Wang, Zhenyuan Chen +9

Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these mo…

cs.MA2026

Living-Harness Is an Interactive-Agent Evolver

Yuetian Du, Yucheng Wang, He Xu +9

Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedba…

cs.CV2026

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

Yuetian Du, Yucheng Wang, Ming Kong +4

Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual accuracy, potentially leading to misdiagno…

cs.CV2026

Linking Perception, Confidence and Accuracy in MLLMs

Yuetian Du, Yucheng Wang, Rongyu Zhang +5

Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual perception to improve accuracy. However, a critical question remains une…

cs.AI2026

Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs

Xiang Zheng, Weiqi Zhai, Wei Wang +15

Recent large language models (LLMs) achieve near-saturation accuracy on many established mathematical reasoning benchmarks, raising concerns about their ability to diagnose genuine…