4 papers · 1 filter
BachVid: Training-Free Video Generation with Consistent Background and Character
Han Yan, Xibin Song, Yifu Wang +3
Diffusion Transformers (DiTs) have recently driven significant progress in text-to-video (T2V) generation. However, generating multiple videos with consistent characters and backgr…
AnchorSync: Global Consistency Optimization for Long Video Editing
Zichi Liu, Yinggui Wang, Tao Wei +1
Editing long videos remains a challenging task due to the need for maintaining both global consistency and temporal coherence across thousands of frames. Existing methods often suf…
PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image
Han Yan, Mingrui Zhang, Yang Li +2
We present PhyCAGE, the first approach for physically plausible compositional 3D asset generation from a single image. Given an input image, we first generate consistent multi-view…
Frankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-Plane
Han Yan, Yang Li, Zhennan Wu +9
We present Frankenstein, a diffusion-based framework that can generate semantic-compositional 3D scenes in a single pass. Unlike existing methods that output a single, unified 3D s…