9 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.CV2023
D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation
Hanjun Li, Xiujun Shu, Sunan He +5
Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a lar…
cs.CV2023★ 3 cited
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding
Taolin Zhang, Sunan He, Dai Tao +3
In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improve…
cs.CV2022★ 9 cited
VLMAE: Vision-Language Masked Autoencoder
Sunan He, Taian Guo, Tao Dai +4
Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…