11 papers
DiT-Reward: Generative Representations for Text-to-Image Reward Modeling
Yuanming Yang, Guoqing Ma, Bo Wang +5
Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative repres…
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predomin…
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Lin Song, Wenbo Li, Guoqing Ma +16
We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatia…
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
Yanbing Zhang, Bo Wang, Jianhui Liu +9
Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static obse…
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7
Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual questio…
TextLDM: Language Modeling with Continuous Latent Diffusion
Jiaxiu Jiang, Jingjing Ren, Wenbo Li +10
Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architect…