activity
20192024
most citedDeepViT: Towards Deeper Vision Transformer

349 citations · 867 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV20245 cited

StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng +2

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents…

cs.CV20243 cited

Magic-Me: Identity-Specific Video Customized Diffusion

Ze Ma, Daquan Zhou, Chun-Hsiao Yeh +6

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven…

cs.CV2024

Sora Generates Videos with Stunning Geometrical Consistency

Xuanyi Li, Daquan Zhou, Chenxu Zhang +3

The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena…

cs.CV2023

ChatAnything: Facetime Chat with LLM-Enhanced Personas

Yilin Zhao, Xinbin Yuan, Shanghua Gao +4

In this technical report, we target generating anthropomorphized personas for LLM-based characters in an online manner, including visual appearance, personality and tones, with onl…

cs.CV2023

Low-Resolution Self-Attention for Semantic Segmentation

Yu-Huan Wu, Shi-Chen Zhang, Yun Liu +6

Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision tra…

cs.CV2023

MaskDiffusion: Boosting Text-to-Image Consistency with Conditional Mask

Yupeng Zhou, Daquan Zhou, Zuo-Liang Zhu +3

Recent advancements in diffusion models have showcased their impressive capacity to generate visually striking images. Nevertheless, ensuring a close match between the generated im…