2 papers
cs.CV2026
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
Zixuan Li, Haokun Lin, Yicheng Xiao +10
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object coun…
cs.CV2025
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
Gong Jingyu, Tong Kunkun, Chen Zhuoran +5
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this pape…