collaborators

12 papers

cs.CV2026

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang +16

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…

cs.CV2026

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

Guoli Jia, Yisheng Zhang, Haote Hu +11

Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing frameworks operate in a training-free man…

cs.CV2026

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

Ruilin Yao, Bo Zhang, Jirui Huang +18

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-wor…

cs.CV2026

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection

Tianyu Wang, Zhiyuan Ma, Qian Wang +3

Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models still generally struggle with a…

cs.CV2025

Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

Guoli Jia, Junyao Hu, Xinwei Long +5

Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emot…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…