3 papers
cs.CV2025
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
Feng Han, Yang Jiao, Shaoxiang Chen +3
The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, cont…
cs.CV2025
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Jiacheng Zhang, Yang Jiao, Shaoxiang Chen +5
Recently, Multimodal Large Language Models (MLLMs) have made significant progress in the video comprehension field. Despite remarkable content reasoning and instruction following c…
cs.CV2025
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
Jiayu Wang, Yang Jiao, Yue Yu +4
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for…