collaborators

6 papers

cs.CV2026

FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters

Shitong Shao, Yufei Gu, Zeke Xie

The recent advent of powerful video generation models, such as Hunyuan, WanX, Veo3, and Kling, has inaugurated a new era in the field. However, the practical deployment of these mo…

cs.CV2025

LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models

Shuai Wang, Daoan Zhang, Tianyi Bai +3

Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-…

cs.CV2025

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

SiXiang Chen, Jianyu Lai, Jialin Gao +11

Generating aesthetic posters is more challenging than simple design images: it requires not only precise text rendering but also the seamless integration of abstract artistic conte…

cs.CV2025

MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis

Shitong Shao, Hongwei Yi, Hanzhong Guo +5

Recently, open-source video diffusion models (VDMs), such as WanX, Magic141 and HunyuanVideo, have been scaled to over 10 billion parameters. These large-scale VDMs have demonstrat…

cs.CV2025

MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Hongwei Yi, Tian Ye, Shitong Shao +10

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse c…

cs.CV2025

Magic 1-For-1: Generating One Minute Video Clips within One Minute

Hongwei Yi, Shitong Shao, Tian Ye +7

In this technical report, we present Magic 1-For-1 (Magic141), an efficient video generation model with optimized memory consumption and inference latency. The key idea is simple:…