8 papers
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
Yuan Wu, Zongxian Yang, Jiayu Qian +5
Large vision-language models (VLMs) often benefit from chain-of-thought (CoT) prompting in general domains, yet its efficacy in medical vision-language tasks remains underexplored.…
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
Jialin Li, Yuan Wu, Yi Chang
To integrate seamlessly into real-world software engineering, Code Agents must evolve from passive instruction followers into proactive collaborative partners. However, current eva…
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
Jinzhe Li, Gengxu Li, Yi Chang +1
Large language models (LLMs) have witnessed rapid advancements, demonstrating remarkable capabilities. However, a notable vulnerability persists: LLMs often uncritically accept fla…
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
Yue Zhou, Yi Chang, Yuan Wu
Reasoning is a critical capability of multimodal large language models (MLLMs) for solving complex multimodal tasks, and judging the correctness of reasoning steps is crucial for i…
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
Haiqi Yang, Jinzhe Li, Gengxu Li +2
Large Multimodal Models (LMMs) have witnessed remarkable growth, showcasing formidable capabilities in handling intricate multimodal tasks with exceptional performance. Recent rese…
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
Jialin Li, Jinzhe Li, Gengxu Li +2
With the advancement of code generation capabilities in large language models (LLMs), their reliance on input premises has intensified. When users provide inputs containing faulty…