11 papers
Scaling Properties of Text Conditioning in Visual Generation
Zilong Chen, Chaorui Deng, Kunchang Li +2
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of…
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
Junke Wang, Xiao Wang, Jiacheng Pan +16
This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…
Multi-Probe Zero Collision Hash (MPZCH): Mitigating Embedding Collisions and Enhancing Model Freshness in Large-Scale Recommenders
Ziliang Zhao, Bi Xue, Emma Lin +16
Embedding tables are critical components of large-scale recommendation systems, facilitating the efficient mapping of high-cardinality categorical features into dense vector repres…
Context Unrolling in Omni Models
Ceyuan Yang, Zhijie Lin, Yang Zhao +16
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such train…
Understanding and Harnessing Sparsity in Unified Multimodal Models
Shwai He, Chaorui Deng, Ang Li +1
Large multimodal models have achieved remarkable progress in both understanding and generation. Recent efforts pursue unified multimodal models that integrate heterogeneous compone…
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
Chenhui Gou, Zilong Chen, Zeyu Wang +10
This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question -- an ability that has recently emerged in prop…