12 papers
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Tengfei Liu, Yang Shi, Yuran Wang +16
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
Guoli Jia, Yisheng Zhang, Haote Hu +11
Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing frameworks operate in a training-free man…
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
Ruilin Yao, Bo Zhang, Jirui Huang +18
Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-wor…
RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection
Tianyu Wang, Zhiyuan Ma, Qian Wang +3
Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models still generally struggle with a…
Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
Guoli Jia, Junyao Hu, Xinwei Long +5
Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emot…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…