4 papers
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
Haonan Han, Jiancheng Huang, Xiaopeng Sun +7
Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current ev…
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
Jie Zhu, Hanghang Ma, Jia Wang +6
In this work, we introduce Wallaroo, a simple autoregressive baseline that leverages next-token prediction to unify multi-modal understanding, image generation, and editing at the…
Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models
Yanbing Zeng, Jia Wang, Hanghang Ma +4
Integrating image generation and understanding into a single framework has become a pivotal goal in the multimodal domain. However, how understanding can effectively assist generat…
LongCat-Image Technical Report
Meituan LongCat Team, Hanghang Ma, Haoxian Tan +10
We introduce LongCat-Image, a pioneering open-source and bilingual (Chinese-English) foundation model for image generation, designed to address core challenges in multilingual text…