130 citations · 151 across the 6 of their papers we have counts for
14 papers · 1 filter
Module-wise Adaptive Distillation for Multimodality Foundation Models
Chen Liang, Jiahui Yu, Ming-Hsuan Yang +5
Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reduci…
MovieCLIP: Visual Scene Recognition in Movies
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli +5
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual sce…
Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models
Rui Qian, Yeqing Li, Zheng Xu +3
Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we ex…
Exploring Temporal Granularity in Self-Supervised Video Representation Learning
Rui Qian, Yeqing Li, Liangzhe Yuan +7
This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and…
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
Golnaz Ghiasi, Yin Cui, Aravind Srinivas +5
Building instance segmentation models that are data-efficient and can handle rare object categories is an important challenge in computer vision. Leveraging data augmentations is a…
Efficient Scale-Permuted Backbone with Learned Resource Distribution
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin +4
Recently, SpineNet has demonstrated promising results on object detection and image classification over ResNet model. However, it is unclear if the improvement adds up when combini…