collaborators

9 papers

cs.CL2025

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?

Zhitao He, Zongwei Lyu, Dazhong Chen +2

Numerous theorems, such as those in geometry, are often presented in multimodal forms (e.g., diagrams). Humans benefit from visual reasoning in such settings, using diagrams to gai…

cs.CV2025

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

Zhiyuan Fan, Yumeng Wang, Sandeep Polisetty +1

Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in…

cs.AI2025

Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

Dadi Guo, Jiayu Liu, Zhiyuan Fan +5

Large reasoning models (e.g., R1, o3) have demonstrated remarkable mathematical problem-solving abilities. However, the high reported accuracy of these advanced models on popular d…

cs.CV2025

MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing

Minghao Liu, Zhitao He, Zhiyuan Fan +2

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation framework…

cs.CV2025

TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models

Jaewoo Lee, Keyang Xuan, Chanakya Ekbote +3

Multimodal Large Language Models (MLLMs) have shown remarkable versatility in understanding diverse multimodal data and tasks. However, these capabilities come with an increased mo…

cs.CL2025

Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization

Shujin Wu, Cheng Qian, Yi R. Fung +2

The growing capabilities of large language models (LLMs) present a key challenge of maintaining effective human oversight. Weak-to-strong generalization (W2SG) offers a promising f…