3 papers
cs.CV2026
DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
Shiyan Du, Conghan Yue, Xinyu Cheng +1
Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic underst…
cs.CV2026
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
Kelaiti Xiao, Liang Yang, Dongyu Zhang +2
We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration o…
cs.CL2025
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
Kelaiti Xiao, Liang Yang, Dongyu Zhang +2
We study idiom-based visual puns--images that align an idiom's literal and figurative meanings--and present an iterative framework that coordinates a large language model (LLM), a…