9 citations · 18 across the 5 of their papers we have counts for
5 papers
See Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval
Xiujun Shu, Wei Wen, Haoqian Wu +5
Text-based person retrieval aims to find the query person based on a textual description. The key is to learn a common latent space mapping between visual-textual modalities. To ac…
TaCo: Textual Attribute Recognition via Contrastive Learning
Chang Nie, Yiqing Hu, Yanqiu Qu +3
As textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing ap…
VLMAE: Vision-Language Masked Autoencoder
Sunan He, Taian Guo, Tao Dai +4
Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…
GMN: Generative Multi-modal Network for Practical Document Information Extraction
Haoyu Cao, Jiefeng Ma, Antai Guo +5
Document Information Extraction (DIE) has attracted increasing attention due to its various advanced applications in the real world. Although recent literature has already achieved…
OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification
Ye Liu, Lingfeng Qiao, Di Yin +4
Scene segmentation and classification (SSC) serve as a critical step towards the field of video structuring analysis. Intuitively, jointly learning of these two tasks can promote e…