7 papers
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
Ruilin Yao, Bo Zhang, Jirui Huang +18
Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-wor…
LAMM-ViT: AI Face Detection via Layer-Aware Modulation of Region-Guided Attention
Jiangling Zhang, Weijie Zhu, Jirui Huang +1
Detecting AI-synthetic faces presents a critical challenge: it is hard to capture consistent structural relationships between facial regions across diverse generation techniques. C…
Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification
Jiangling Zhang, Shuxuan Gao, Bofan Liu +4
The proliferation of highly realistic AI-generated images poses critical challenges for digital forensics, demanding precise pixel-level localization of manipulated regions. Existi…
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
Shengwu. Xiong, Tianyu. Zou, Cong. Wang +1
Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristi…
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
Jinghao Huang, Yaxiong Chen, Ganchao Liu
With the advancement of drone technology, the volume of video data increases rapidly, creating an urgent need for efficient semantic retrieval. We are the first to systematically p…
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
Shilan Zhang, Jirui Huang, Ruilin Yao +4
Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through…