4 papers · 1 filter
Wan-R1: Verifiable-Reinforcement Learning for Video Reasoning
Ming Liu, Yunbei Zhang, Shilong Liu +2
Video generation models produce visually coherent content but struggle with tasks requiring spatial reasoning and multi-step planning. Reinforcement learning (RL) offers a path to…
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
Ming Liu, Siyuan Liang, Koushik Howlader +3
Vision-Language Models (VLMs) have been integrated into autonomous driving systems to enhance reasoning capabilities through tasks such as Visual Question Answering (VQA). However,…
Is Your Video Language Model a Reliable Judge?
Ming Liu, Wensheng Zhang
As video language models (VLMs) gain more applications in various scenarios, the need for robust and scalable evaluation of their performance becomes increasingly critical. The tra…
On the robustness of multimodal language model towards distractions
Ming Liu, Hao Chen, Jindong Wang +1
Although vision-language models (VLMs) have achieved significant success in various applications such as visual question answering, their resilience to prompt variations remains an…