2 citations · 5 across the 5 of their papers we have counts for
5 papers
Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning
Yeming Chen, Siyu Zhang, Yaoru Sun +2
With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…
S3IM: Stochastic Structural SIMilarity and Its Unreasonable Effectiveness for Neural Fields
Zeke Xie, Xindi Yang, Yujie Yang +5
Recently, Neural Radiance Field (NeRF) has shown great success in rendering novel-view images of a given scene by learning an implicit representation with only posed RGB images. Ne…
ChatGPT-Crawler: Find out if ChatGPT really knows what it's talking about
Aman Rangapur, Haoran Wang
Large language models have gained considerable interest for their impressive performance on various tasks. Among these models, ChatGPT developed by OpenAI has become extremely popu…
Boosting Video-Text Retrieval with Explicit High-Level Semantics
Haoran Wang, Di Xu, Dongliang He +4
Video-text retrieval (VTR) is an attractive yet challenging task for multi-modal understanding, which aims to search for relevant video (text) given a query (video). Existing metho…
NSNet: Non-saliency Suppression Sampler for Efficient Video Recognition
Boyang Xia, Wenhao Wu, Haoran Wang +5
It is challenging for artificial intelligence systems to achieve accurate video recognition under the scenario of low computation costs. Adaptive inference based efficient video re…