13 papers
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
Yinan Zhou, Haokun Lin, Yichen Wu +7
Large multimodal models have achieved strong reasoning on complex visual tasks, but their inference efficiency is often restricted by long chains of thought. A promising solution i…
StepGuard: Guarding Web Navigation via Single-Step Calibration
Zhihao Cui, Yuchen Zhang, Xiyang Sun +6
Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and re…
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
Yuchen Zhang, Yaxiong Wang, Kecheng Han +4
Recent advances in generative AI have significantly enhanced the realism of multimodal media manipulation, thereby posing substantial challenges to manipulation detection. Existing…
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
Jingchun Lian, Lingyu Liu, Yaxiong Wang +4
Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulatio…
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
Lingyu Liu, Yaxiong Wang, Li Zhu +1
Video frame interpolation aims to synthesize realistic intermediate frames between given endpoints while adhering to specific motion semantics. While recent generative models have…
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
Shuyu Yang, Yaxiong Wang, Yongrui Li +2
In this work, we focus on text-based person retrieval, which identifies individuals based on textual descriptions. Despite advancements enabled by synthetic data for pretraining, a…