1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Zixuan Li, Haokun Lin, Yicheng Xiao +10
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object coun…