21 citations · 25 across the 5 of their papers we have counts for
4 papers · 1 filter
HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion Model
Yi Wang, Jian Ma, Ruizhi Shao +3
This paper aims to generate physically-layered 3D humans from text prompts. Existing methods either generate 3D clothed humans as a whole or support only tight and simple clothing…
AUG: A New Dataset and An Efficient Model for Aerial Image Urban Scene Graph Generation
Yansheng Li, Kun Li, Yongjun Zhang +2
Scene graph generation (SGG) aims to understand the visual objects and their semantic relationships from one given image. Until now, lots of SGG datasets with the eyelevel view are…
Data Augmentation for Human Behavior Analysis in Multi-Person Conversations
Kun Li, Dan Guo, Guoliang Chen +2
In this paper, we present the solution of our team HFUT-VUT for the MultiMediate Grand Challenge 2023 at ACM Multimedia 2023. The solution covers three sub-challenges: bodily behav…
HRVQA: A Visual Question Answering Benchmark for High-Resolution Aerial Images
Kun Li, George Vosselman, Michael Ying Yang
Visual question answering (VQA) is an important and challenging multimodal task in computer vision. Recently, a few efforts have been made to bring VQA task to aerial images, due t…