11 papers
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Tianyi Gao, Han Fang, Tianyi Ding +9
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…
VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation
Songyu Xu, Xin Wang, Qiang Chen +6
The paper introduces VGIF-Score, an automated and interpretable framework that evaluates how well video generation models follow long, compositional instructions by parsing prompts…
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
Yuanzhi Liu, Shousheng Zhao, Bo Zhou +2
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data conta…
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
Baoteng Li, Xianghao Zang, Xinran Wang +8
Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based on Group Relative Policy Optimi…
DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents
Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7
Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
Xinran Wang, Yuxuan Zhang, Xiao Zhang +7
Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), capti…