activity
20212026
most citedWhisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers

61 citations · 196 across the 42 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

5 papers · 2 filters

cs.CV2023★ 1 cited

Learning Human Action Recognition Representations Without Real Humans

Howard Zhong, Samarth Mishra, Donghyun Kim +7

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets…

cs.CV2023★ 12 cited

Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Sivan Doveh, Assaf Arbelle, Sivan Harary +9

Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text, leading to numerous applications such as cross-modal retrieval, vis…

cs.CV2023

Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs

Roei Herzig, Alon Mendelson, Leonid Karlinsky +4

Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks. However, recent works have shown that even the best VLMs struggle t…

cs.CV2023

Going Beyond Nouns With Vision & Language Models Using Synthetic Data

Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith +8

Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot ope…

cs.CV2023★ 2 cited

MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Wei Lin, Leonid Karlinsky, Nina Shvetsova +6

Large scale Vision-Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot…