162 citations · 402 across the 48 of their papers we have counts for
55 papers · 1 filter
DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
Shengming Yin, Chenfei Wu, Jian Liang +4
Controllable video generation has gained significant attention in recent years. However, two main limitations persist: Firstly, most existing works focus on either text, image, or…
Detect Any Shadow: Segment Anything for Video Shadow Detection
Yonghui Wang, Wengang Zhou, Yunyao Mao +1
Segment anything model (SAM) has achieved great success in the field of natural image segmentation. Nevertheless, SAM tends to consider shadows as background and therefore does not…
Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding
Yuechen Wang, Wengang Zhou, Houqiang Li
Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotati…
UDoc-GAN: Unpaired Document Illumination Correction with Background Light Prior
Yonghui Wang, Wengang Zhou, Zhenbo Lu +1
Document images captured by mobile devices are usually degraded by uncontrollable illumination, which hampers the clarity of document content. Recently, a series of research effort…
Multi-Target Active Object Tracking with Monte Carlo Tree Search and Target Motion Modeling
Zheng Chen, Jian Zhao, Mingyu Yang +2
In this work, we are dedicated to multi-target active object tracking (AOT), where there are multiple targets as well as multiple cameras in the environment. The goal is maximize t…
Learning Enriched Illuminants for Cross and Single Sensor Color Constancy
Xiaodong Cun, Zhendong Wang, Chi-Man Pun +4
Color constancy aims to restore the constant colors of a scene under different illuminants. However, due to the existence of camera spectral sensitivity, the network trained on a c…