7 papers
Instruction-Based Video Editing by Repurposing an Image Editing Model
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi +1
Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source…
DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi
Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing a…
PE-Field 4D: Video Generation Models as Canvas
Yunpeng Bai, Haoxiang Li, Qixing Huang
Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging.…
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
Shaoheng Fang, Hanwen Jiang, Yunpeng Bai +2
Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally co…
GeoVideo: Introducing Geometric Regularization into Video Generation Model
Yunpeng Bai, Shaoheng Fang, Chaohui Yu +2
Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches op…
Positional Encoding Field
Yunpeng Bai, Haoxiang Li, Qixing Huang
Diffusion Transformers (DiTs) have emerged as the dominant architecture for visual generation, powering state-of-the-art image and video models. By representing images as patch tok…