11 papers
Scaling Properties of Text Conditioning in Visual Generation
Zilong Chen, Chaorui Deng, Kunchang Li +2
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of…
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Ruofei Ju, Xinrui Wang, Xin Ding +12
Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied environments vary across layouts,…
MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation
Dewei Zhou, Xinyu Huang, Xun Wang +6
Generative visual models fundamentally struggle with precise spatial control. This arises from a core disconnect: models can process textual descriptions of space but cannot direct…
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
Yabo Zhang, Kunchang Li, Dewei Zhou +2
While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance un…
Context Unrolling in Omni Models
Ceyuan Yang, Zhijie Lin, Yang Zhao +16
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such train…
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
Jie Liu, Zilyu Ye, Linxiao Yuan +8
Unified models capable of interleaved generation have emerged as a promising paradigm, with the community increasingly converging on autoregressive modeling for text and flow match…