56 citations · 61 across the 6 of their papers we have counts for
6 papers
NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
Shengming Yin, Chenfei Wu, Huan Yang +13
In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…
Learning 3D Photography Videos via Self-supervised Diffusion on Single Images
Xiaodong Wang, Chenfei Wu, Shengming Yin +9
3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input f…
A Unified Multi-view Multi-person Tracking Framework
Fan Yang, Shigeyuki Odashima, Sosuke Yamao +3
Although there is a significant development in 3D Multi-view Multi-person Tracking (3D MM-Tracking), current 3D MM-Tracking frameworks are designed separately for footprint and pos…
A Multi-Person Video Dataset Annotation Method of Spatio-Temporally Actions
Fan Yang
Spatio-temporal action detection is an important and challenging problem in video understanding. However, the application of the existing large-scale spatio-temporal action dataset…
A Hierarchical Mixture Density Network
Fan Yang, Jaymar Soriano, Takatomi Kubo +1
The relationship among three correlated variables could be very sophisticated, as a result, we may not be able to find their hidden causality and model their relationship explicitl…
Data Augmentation for Object Detection via Progressive and Selective Instance-Switching
Hao Wang, Qilong Wang, Fan Yang +2
Collection of massive well-annotated samples is effective in improving object detection performance but is extremely laborious and costly. Instead of data collection and annotation…