4 papers · 1 filter
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
Feng Han, Yang Jiao, Shaoxiang Chen +3
The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, cont…
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Jiacheng Zhang, Yang Jiao, Shaoxiang Chen +5
Recently, Multimodal Large Language Models (MLLMs) have made significant progress in the video comprehension field. Despite remarkable content reasoning and instruction following c…
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
Jiayu Wang, Yang Jiao, Yue Yu +4
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for…
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
Jiacheng Zhang, Yang Jiao, Shaoxiang Chen +2
Recently, Multimodal Large Language Models (MLLMs) have sparked great research interests owing to their exceptional content-reasoning and instruction-following capabilities. To eff…