14 papers
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
Junhao Liu, Jian-Wei Zhang, Tao Huang +3
Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial inst…
Rosetta: Composable Native Multimodal Pretraining
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuou…
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts. While existing parad…
TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing
Peizhen Zhang, Yang Li, Xunsong Li +10
Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their enormous parameter consumption…
AdaMem: Learning What to Remember for Personalized Long-Horizon LLM Agents
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Long-term memory systems for Large Language Model (LLM) agents typically try to \emph{remember everything}, extracting memories uniformly to retain as many facts as possible. In pr…
CrossFlow: One-Step Generation Across Latent and Pixel Spaces
Xiyuan Wang, Xiao Zhang, Yang Li +4
Most diffusion and flow-matching generators define the prior, probability path, and prediction target in the same representation space. Latent diffusion improves efficiency by movi…