26 citations · 74 across the 12 of their papers we have counts for
9 papers · 1 filter
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Arun Reddy, Alexander Martin, Eugene Yang +7
In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval…
Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks
Micah Goldblum, Hossein Souri, Renkun Ni +10
Neural network based computer vision systems are typically built on a backbone, a pretrained or randomly initialized feature extractor. Several years ago, the default option was an…
Whole-body Detection, Recognition and Identification at Altitude and Range
Siyuan Huang, Ram Prabhakar Kathirvel, Chun Pong Lau +1
In this paper, we address the challenging task of whole-body biometric detection, recognition, and identification at distances of up to 500m and large pitch angles of up to 50 degr…
EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
Shraman Pramanick, Yale Song, Sayan Nag +5
Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP framewor…
SMC-UDA: Structure-Modal Constraint for Unsupervised Cross-Domain Renal Segmentation
Zhusi Zhong, Jie Li, Lulu Bi +6
Medical image segmentation based on deep learning often fails when deployed on images from a different domain. The domain adaptation methods aim to solve domain-shift challenges, b…
Self-Denoising Neural Networks for Few Shot Learning
Steven Schwarcz, Sai Saketh Rambhatla, Rama Chellappa
In this paper, we introduce a new architecture for few shot learning, the task of teaching a neural network from as few as one or five labeled examples. Inspired by the theoretical…