1 citations · 1 across the 5 of their papers we have counts for
5 papers
Vision as Unified Multimodal Generation
Xiaoyang Han, Jianhua Li, Kewang Deng +14
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal…
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng +55
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…
Scaling Spatial Intelligence with Multimodal Foundation Models
Zhongang Cai, Ruisi Wang, Chenyang Gu +26
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation m…
ConsistCompose: Unified Multimodal Layout Control for Image Composition
Xuanke Shi, Boxuan Li, Xiaoyang Han +4
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with imag…
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
Zhongang Cai, Yubo Wang, Qingping Sun +21
Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very cap…