activity
20232026
most citedMicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Xingjian Leng, Jaskirat Singh, Zhanhao Liang +5

Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods improve the FID and rel…

cs.CV2026

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

Zhicong Tang, Zhao Zhang, Jingye Chen +6

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editin…

cs.CV2024

Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Zeyu Liu, Weicong Liang, Zhanhao Liang +4

Visual text rendering poses a fundamental challenge for contemporary text-to-image generation models, with the core problem lying in text encoder deficiencies. To achieve accurate…

cs.CV20231 cited

MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation

Yanhui Wang, Jianmin Bao, Wenming Weng +12

We present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with vi…

cs.CV2023

ARTV: Auto-Regressive Text-to-Video Generation with Diffusion Models

Wenming Weng, Ruoyu Feng, Yanhui Wang +10

We present ARTV, an efficient framework for auto-regressive video generation with diffusion models. Unlike existing methods that generate entire videos in one-s…