92 citations · 154 across the 6 of their papers we have counts for
9 papers · 1 filter
Beyond Language Modeling: An Exploration of Multimodal Pretraining
Shengbang Tong, David Fan, John Nguyen +18
The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models r…
Distilling Self-Supervised Vision Transformers for Weakly-Supervised Few-Shot Classification & Segmentation
Dahyun Kang, Piotr Koniusz, Minsu Cho +1
We address the task of weakly-supervised few-shot image classification and segmentation, by leveraging a Vision Transformer (ViT) pretrained with self-supervision. Our proposed met…
Time-rEversed diffusioN tEnsor Transformer: A new TENET of Few-Shot Object Detection
Shan Zhang, Naila Murray, Lei Wang +1
In this paper, we tackle the challenging problem of Few-shot Object Detection. Existing FSOD pipelines (i) use average-pooled representations that result in information loss; and/o…
Temporally-Weighted Hierarchical Clustering for Unsupervised Action Segmentation
M. Saquib Sarfraz, Naila Murray, Vivek Sharma +3
Action segmentation refers to inferring boundaries of semantically consistent visual concepts in videos and is an important requirement for many video understanding tasks. For this…
Virtual KITTI 2
Yohann Cabon, Naila Murray, Martin Humenberger
This paper introduces an updated version of the well-known Virtual KITTI dataset which consists of 5 sequence clones from the KITTI tracking benchmark. In addition, the dataset pro…
Generating Human Action Videos by Coupling 3D Game Engines and Probabilistic Graphical Models
César Roberto de Souza, Adrien Gaidon, Yohann Cabon +2
Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtai…