5 papers
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
Yibo Jiang, Tao Wu, Rui Jiang +4
Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where…
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
Yuyuan Yang, Junkun Hong, Hongrong Wang +11
Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…
SpatialText: A Pure-Text Cognitive Benchmark for Spatial Understanding in Large Language Models
Peiyao Jiang, Zequn Qin, Xi Li
Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely pr…
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
Tao Wu, Yibo Jiang, Yehao Lu +4
Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human a…
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
Longrong Yang, Zhixiong Zeng, Yufeng Zhong +7
Multimodal large language models are evolving toward multimodal agents capable of proactively executing tasks. Most agent research focuses on GUI or embodied scenarios, which corre…