7 papers
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems
Quanxing Xu, Yuhao Tian, Ling Zhou +4
Visual Question Answering (VQA), as the representative multimodal task, serves as a key benchmark for evaluating the reasoning capabilities of Multimodal Large Language Models (MLL…
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
Quanxing Xu, Ling Zhou, Xian Zhong +3
With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of multimodal tasks. As a core…
SCP: Spatial Causal Prediction in Video
Yanguang Zhao, Jie Yang, Shengqiong Wu +9
Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as…
Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
Wenxuan Liu, Zhuo Zhou, Xuemei Jia +4
Action recognition in unmanned aerial vehicles (UAVs) poses unique challenges due to significant view variations along the vertical spatial axis. Unlike traditional ground-based se…
OccludeNet: A Causal Journey into Mixed-View Actor-Centric Video Action Recognition under Occlusions
Guanyu Zhou, Wenxuan Liu, Wenxin Huang +3
The lack of occlusion data in common action recognition video datasets limits model robustness and hinders consistent performance gains. We build OccludeNet, a large-scale occluded…
QIRL: Optimized Question-Image Relation Learning for Bias-Robust Visual Question Answering
Quanxing Xu, Ling Zhou, Xian Zhong +3
Existing bias mitigation methods for Visual Question Answering (VQA), a typical Artificial intelligence application, endure two main limitations. First, they fail to capture the op…