activity
20222024
most citedModelScope Text-to-Video Technical Report

47 citations · 127 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV202434 cited

Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification

Wenjia Xu, Jiuniu Wang, Zhiwei Wei +2

Deep neural networks have achieved promising progress in remote sensing (RS) image classification, for which the training process requires abundant samples for each class. However,…

cs.CV20233 cited

Jointly Optimized Global-Local Visual Localization of UAVs

Haoling Li, Jiuniu Wang, Zhiwei Wei +1

Navigation and localization of UAVs present a challenge when global navigation satellite systems (GNSS) are disrupted and unreliable. Traditional techniques, such as simultaneous l…

cs.CV202347 cited

ModelScope Text-to-Video Technical Report

Jiuniu Wang, Hangjie Yuan, Dayou Chen +3

This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-t…

cs.CV202343 cited

VideoComposer: Compositional Video Synthesis with Motion Controllability

Xiang Wang, Hangjie Yuan, Shiwei Zhang +6

The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable vide…

cs.CV2022

Distinctive Image Captioning via CLIP Guided Group Optimization

Youyuan Zhang, Jiuniu Wang, Hao Wu +1

Image captioning models are usually trained according to human annotated ground-truth captions, which could generate accurate but generic captions. In this paper, we focus on gener…

cs.CV2022

Multi-dimension Geospatial feature learning for urban region function recognition

Wenjia Xu, Jiuniu Wang, Yirong Wu

Urban region function recognition plays a vital character in monitoring and managing the limited urban areas. Since urban functions are complex and full of social-economic properti…