4 papers
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
Zhihao Zhu, Jiafeng Liang, Shixin Jiang +5
Large Multimodal Models (LMMs) have demonstrated impressive capabilities in video reasoning via Chain-of-Thought (CoT). However, the robustness of their reasoning chains remains qu…
Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining
Zhaocheng Yu, Kui Jiang, Junjun Jiang +3
Rain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining method…
Leveraging Static Relationships for Intra-Type and Inter-Type Message Passing in Video Question Answering
Lili Liang, Guanglu Sun
Video Question Answering (VideoQA) is an important research direction in the field of artificial intelligence, enabling machines to understand video content and perform reasoning a…
Unbiased Scene Graph Generation by Type-Aware Message Passing on Heterogeneous and Dual Graphs
Guanglu Sun, Jin Qiu, Lili Liang
Although great progress has been made in the research of unbiased scene graph generation, issues still hinder improving the predictive performance of both head and tail classes. An…