activity
20222024
most citedDual-Flattening Transformers through Decomposed Row and Column Queries for Semantic Segmentation

3 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling

Zilyu Ye, Jinxiu Liu, Ruotian Peng +9

Recent image generation models excel at creating high-quality images from brief captions. However, they fail to maintain consistency of multiple instances across images when encoun…

cs.CV2024

UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures

Mingyuan Zhou, Rakib Hyder, Ziwei Xuan +1

Recent advances in 3D avatar generation have gained significant attentions. These breakthroughs aim to produce more realistic animatable avatars, narrowing the gap between virtual…

cs.CV2023

OmniMotionGPT: Animal Motion Generation with Limited Data

Zhangsihao Yang, Mingyuan Zhou, Mengyi Shan +6

Our paper aims to generate diverse and realistic animal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven…

cs.CV2023

AdPE: Adversarial Positional Embeddings for Pretraining Vision Transformers via MAE+

Xiao Wang, Ying Wang, Ziwei Xuan +1

Unsupervised learning of vision transformers seeks to pretrain an encoder via pretext tasks without labels. Among them is the Masked Image Modeling (MIM) aligned with pretraining o…

cs.CV2022★ 3 cited

Dual-Flattening Transformers through Decomposed Row and Column Queries for Semantic Segmentation

Ying Wang, Chiuman Ho, Wenju Xu +3

It is critical to obtain high resolution features with long range dependency for dense prediction tasks such as semantic segmentation. To generate high-resolution output of size $H…