162 citations · 393 across the 47 of their papers we have counts for
64 papers
Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding
Yuechen Wang, Wengang Zhou, Houqiang Li
Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotati…
UDoc-GAN: Unpaired Document Illumination Correction with Background Light Prior
Yonghui Wang, Wengang Zhou, Zhenbo Lu +1
Document images captured by mobile devices are usually degraded by uncontrollable illumination, which hampers the clarity of document content. Recently, a series of research effort…
Multi-Target Active Object Tracking with Monte Carlo Tree Search and Target Motion Modeling
Zheng Chen, Jian Zhao, Mingyu Yang +2
In this work, we are dedicated to multi-target active object tracking (AOT), where there are multiple targets as well as multiple cameras in the environment. The goal is maximize t…
Learning Enriched Illuminants for Cross and Single Sensor Color Constancy
Xiaodong Cun, Zhendong Wang, Chi-Man Pun +4
Color constancy aims to restore the constant colors of a scene under different illuminants. However, due to the existence of camera spectral sensitivity, the network trained on a c…
MVP: Multimodality-guided Visual Pre-training
Longhui Wei, Lingxi Xie, Wengang Zhou +2
Recently, masked image modeling (MIM) has become a promising direction for visual pre-training. In the context of vision transformers, MIM learns effective visual representation by…
Coordinate-Aligned Multi-Camera Collaboration for Active Multi-Object Tracking
Zeyu Fang, Jian Zhao, Mingyu Yang +3
Active Multi-Object Tracking (AMOT) is a task where cameras are controlled by a centralized system to adjust their poses automatically and collaboratively so as to maximize the cov…