11 citations · 11 across the 4 of their papers we have counts for
7 papers · 1 filter
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
Hao Yang, Zhiyu Tan, Jia Gong +7
We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified vi…
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
Mengping Yang, Zhiyu Tan, Binglei Li +3
Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturi…
SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models
Hesen Chen, Junyan Wang, Zhiyu Tan +1
Modern diffusion models encounter a fundamental trade-off between training efficiency and generation quality. While existing representation alignment methods, such as REPA, acceler…
Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos
Zhiyu Tan, Junyan Wang, Hao Yang +4
Text-to-video generation has demonstrated promising progress with the advent of diffusion models, yet existing approaches are limited by dataset quality and computational resources…
E2ED^2:Direct Mapping from Noise to Data for Enhanced Diffusion Models
Zhiyu Tan, WenXu Qian, Hesen Chen +3
Diffusion models have established themselves as the de facto primary paradigm in visual generative modeling, revolutionizing the field through remarkable success across various div…
Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition
Ming Lin, Pichao Wang, Zhenhong Sun +5
Accuracy predictor is a key component in Neural Architecture Search (NAS) for ranking architectures. Building a high-quality accuracy predictor usually costs enormous computation.…