most citedContextual Modeling for 3D Dense Captioning on Point Clouds

8 citations · 14 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2022

A Unified Framework for Contrastive Learning from a Perspective of Affinity Matrix

Wenbin Li, Meihao Kong, Xuesong Yang +4

In recent years, a variety of contrastive learning based unsupervised visual representation learning methods have been designed and achieved great success in many visual tasks. Gen…

cs.CV20228 cited

Contextual Modeling for 3D Dense Captioning on Point Clouds

Yufeng Zhong, Long Xu, Jiebo Luo +1

3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence fo…

cs.CV2021

LUAI Challenge 2021 on Learning to Understand Aerial Images

Gui-Song Xia, Jian Ding, Ming Qian +33

This report summarizes the results of Learning to Understand Aerial Images (LUAI) 2021 challenge held on ICCV 2021, which focuses on object detection and semantic segmentation in a…

cs.CV20213 cited

Multi-Modulation Network for Audio-Visual Event Localization

Hao Wang, Zheng-Jun Zha, Liang Li +2

We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the…

cs.CV20213 cited

Memory Enhanced Embedding Learning for Cross-Modal Video-Text Retrieval

Rui Zhao, Kecheng Zheng, Zheng-Jun Zha +2

Cross-modal video-text retrieval, a challenging task in the field of vision and language, aims at retrieving corresponding instance giving sample from either modality. Existing app…