most citedControlVideo: Training-free Controllable Text-to-Video Generation

34 citations · 49 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2024

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

Yabo Zhang, Yuxiang Wei, Xianhui Lin +5

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V)…

cs.CV20233 cited

VQ-Font: Few-Shot Font Generation with Structure-Aware Enhancement and Quantization

Mingshuai Yao, Yabo Zhang, Xianhui Lin +2

Few-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which…

cs.CV202334 cited

ControlVideo: Training-free Controllable Text-to-Video Generation

Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temp…

cs.CV20231 cited

Associating Spatially-Consistent Grouping with Text-supervised Semantic Segmentation

Yabo Zhang, Zihao Wang, Jun Hao Liew +4

In this work, we investigate performing semantic segmentation solely through the training on image-sentence pairs. Due to the lack of dense annotations, existing text-supervised me…

cs.CV202211 cited

Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks

Yabo Zhang, Mingshuai Yao, Yuxiang Wei +3

One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for…