9 citations · 19 across the 25 of their papers we have counts for
Showing 2024 · cs.CVShow all
3 papers · 2 filters
cs.CV2024★ 4 cited
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
Bo Cheng, Yuhang Ma, Liebucha Wu +5
The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout…
cs.CV2024
Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task
Jing Wang, Ao Ma, Jiasong Feng +3
The global self-attention mechanism in diffusion transformers involves redundant computation due to the sparse and redundant nature of visual information, and the attention map of…
cs.CV2024★ 1 cited
FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance
Jiasong Feng, Ao Ma, Jing Wang +2
Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T…