5 papers
SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
Ruoyu Feng, Jinming Liu, Yuqi Wang +7
Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate the training process, yet th…
Bridging Video Understanding and Generation in a Unified Framework
Yuqi Wang, Runyi Li, Ruoyu Feng +3
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexp…
GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
Yueming Pan, Ruoyu Feng, Jianmin Bao +2
Video outpainting extends a video beyond its original boundaries by synthesizing missing border content. Compared with image outpainting, it requires not only per-frame spatial pla…
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
Yueming Pan, Ruoyu Feng, Qi Dai +5
Latent Diffusion Models (LDMs) inherently follow a coarse-to-fine generation process, where high-level semantic structure is generated slightly earlier than fine-grained texture. T…
ContentV: Efficient Training of Video Generation Models with Limited Compute
Wenfeng Lin, Renjie Chen, Boyuan Liu +10
Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter…