collaborators

8 papers

cs.CV2026

Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine

Yuan Wu, Zongxian Yang, Jiayu Qian +5

Large vision-language models (VLMs) often benefit from chain-of-thought (CoT) prompting in general domains, yet its efficacy in medical vision-language tasks remains underexplored.…

cs.SE2026

ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions

Jialin Li, Yuan Wu, Yi Chang

To integrate seamlessly into real-world software engineering, Code Agents must evolve from passive instruction followers into proactive collaborative partners. However, current eva…

cs.CL2025

Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models

Jinzhe Li, Gengxu Li, Yi Chang +1

Large language models (LLMs) have witnessed rapid advancements, demonstrating remarkable capabilities. However, a notable vulnerability persists: LLMs often uncritically accept fla…

cs.AI2025

ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges

Yue Zhou, Yi Chang, Yuan Wu

Reasoning is a critical capability of multimodal large language models (MLLMs) for solving complex multimodal tasks, and judging the correctness of reasoning steps is crucial for i…

cs.CV2025

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability

Haiqi Yang, Jinzhe Li, Gengxu Li +2

Large Multimodal Models (LMMs) have witnessed remarkable growth, showcasing formidable capabilities in handling intricate multimodal tasks with exceptional performance. Recent rese…

cs.AI2025

Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework

Jialin Li, Jinzhe Li, Gengxu Li +2

With the advancement of code generation capabilities in large language models (LLMs), their reliance on input premises has intensified. When users provide inputs containing faulty…