most citedSee Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval

9 citations · 18 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023

D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation

Hanjun Li, Xiujun Shu, Sunan He +5

Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a lar…

cs.CV2023

Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies

Bei Gan, Xiujun Shu, Ruizhi Qiao +4

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1…

cs.CV20229 cited

See Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval

Xiujun Shu, Wei Wen, Haoqian Wu +5

Text-based person retrieval aims to find the query person based on a textual description. The key is to learn a common latent space mapping between visual-textual modalities. To ac…

cs.CV20229 cited

VLMAE: Vision-Language Masked Autoencoder

Sunan He, Taian Guo, Tao Dai +4

Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…

cs.CV2022

Exploiting Feature Diversity for Make-up Temporal Video Grounding

Xiujun Shu, Wei Wen, Taian Guo +3

This technical report presents the 3rd winning solution for MTVG, a new task introduced in the 4-th Person in Context (PIC) Challenge at ACM MM 2022. MTVG aims at localizing the te…