8 citations · 14 across the 5 of their papers we have counts for
5 papers
A Unified Framework for Contrastive Learning from a Perspective of Affinity Matrix
Wenbin Li, Meihao Kong, Xuesong Yang +4
In recent years, a variety of contrastive learning based unsupervised visual representation learning methods have been designed and achieved great success in many visual tasks. Gen…
Contextual Modeling for 3D Dense Captioning on Point Clouds
Yufeng Zhong, Long Xu, Jiebo Luo +1
3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence fo…
LUAI Challenge 2021 on Learning to Understand Aerial Images
Gui-Song Xia, Jian Ding, Ming Qian +33
This report summarizes the results of Learning to Understand Aerial Images (LUAI) 2021 challenge held on ICCV 2021, which focuses on object detection and semantic segmentation in a…
Multi-Modulation Network for Audio-Visual Event Localization
Hao Wang, Zheng-Jun Zha, Liang Li +2
We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the…
Memory Enhanced Embedding Learning for Cross-Modal Video-Text Retrieval
Rui Zhao, Kecheng Zheng, Zheng-Jun Zha +2
Cross-modal video-text retrieval, a challenging task in the field of vision and language, aims at retrieving corresponding instance giving sample from either modality. Existing app…