collaborators

6 papers

cs.AI2025

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

Ming Liu, Wensheng Zhang

Multimodal Large Language Models (MLLMs) show impressive vision-language benchmark performance, yet growing concerns about data contamination (test set exposure during training) ri…

cs.CL2025

Is your multimodal large language model a good science tutor?

Ming Liu, Liwen Wang, Wensheng Zhang

Multimodal large language models (MLLMs) demonstrate impressive performance on scientific reasoning tasks (e.g., ScienceQA). However, most existing benchmarks focus narrowly on the…

cs.CV2025

Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving

Ming Liu, Siyuan Liang, Koushik Howlader +3

Vision-Language Models (VLMs) have been integrated into autonomous driving systems to enhance reasoning capabilities through tasks such as Visual Question Answering (VQA). However,…

cs.CV2025

Is Your Video Language Model a Reliable Judge?

Ming Liu, Wensheng Zhang

As video language models (VLMs) gain more applications in various scenarios, the need for robust and scalable evaluation of their performance becomes increasingly critical. The tra…

cs.CV2025

On the robustness of multimodal language model towards distractions

Ming Liu, Hao Chen, Jindong Wang +1

Although vision-language models (VLMs) have achieved significant success in various applications such as visual question answering, their resilience to prompt variations remains an…

cs.CL2025

On Fairness of Unified Multimodal Large Language Model for Image Generation

Ming Liu, Hao Chen, Jindong Wang +3

Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in visual understanding and generation in an end-to-end pipeline. Compared with generati…