79 citations · 215 across the 12 of their papers we have counts for
12 papers · 1 filter
Rotation Invariant Transformer for Recognizing Object in UAVs
Shuoyi Chen, Mang Ye, Bo Du
Recognizing a target of interest from the UAVs is much more challenging than the existing object re-identification tasks across multiple city cameras. The images taken by the UAVs…
An Empirical Study of CLIP for Text-based Person Search
Min Cao, Yang Bai, Ziyin Zeng +2
Text-based Person Search (TBPS) aims to retrieve the person images using natural language descriptions. Recently, Contrastive Language Image Pretraining (CLIP), a universal large c…
Symmetric Uncertainty-Aware Feature Transmission for Depth Super-Resolution
Wuxuan Shi, Mang Ye, Bo Du
Color-guided depth super-resolution (DSR) is an encouraging paradigm that enhances a low-resolution (LR) depth map guided by an extra high-resolution (HR) RGB image from the same s…
Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval
Ding Jiang, Mang Ye
Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual mo…
Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning
Xian Zhong, Zipeng Li, Shuqin Chen +3
Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual repre…
The Multi-Modal Video Reasoning and Analyzing Competition
Haoran Peng, He Huang, Li Xu +15
In this paper, we introduce the Multi-Modal Video Reasoning and Analyzing Competition (MMVRAC) workshop in conjunction with ICCV 2021. This competition is composed of four differen…