35 citations · 41 across the 3 of their papers we have counts for
4 papers
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
Peng Jin, Jinfa Huang, Fenglin Liu +5
Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to t…
Fuzzy Positive Learning for Semi-supervised Semantic Segmentation
Pengchong Qiao, Zhidan Wei, Yu Wang +6
Semi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevita…
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
Mengjun Cheng, Yipeng Sun, Longchao Wang +8
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable…
Harmonized Multimodal Learning with Gaussian Process Latent Variable Models
Guoli Song, Shuhui Wang, Qingming Huang +1
Multimodal learning aims to discover the relationship between multiple modalities. It has become an important research topic due to extensive multimodal applications such as cross-…