most citedMultitask Prompt Tuning Enables Parameter-Efficient Transfer Learning

30 citations · 61 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV20231 cited

Learning Human Action Recognition Representations Without Real Humans

Howard Zhong, Samarth Mishra, Donghyun Kim +7

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets…

cs.LG2023

GeRA: Label-Efficient Geometrically Regularized Alignment

Dustin Klebe, Tal Shnitzer, Mikhail Yurochkin +2

Pretrained unimodal encoders incorporate rich semantic information into embedding space structures. To be similarly informative, multi-modal encoders typically require massive amou…

cs.CV2023

TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification

M. Jehanzeb Mirza, Leonid Karlinsky, Wei Lin +3

Vision and Language Models (VLMs), such as CLIP, have enabled visual recognition of a potentially unlimited set of categories described by text prompts. However, for the best visua…

cs.CV202312 cited

Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Sivan Doveh, Assaf Arbelle, Sivan Harary +9

Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text, leading to numerous applications such as cross-modal retrieval, vis…

cs.CL2023

Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages

Andrew Rouditchenko, Sameer Khurana, Samuel Thomas +6

Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there ar…

cs.CV20231 cited

Constructive Assimilation: Boosting Contrastive Learning Performance through View Generation Strategies

Ligong Han, Seungwook Han, Shivchander Sudalairaj +8

Transformations based on domain expertise (expert transformations), such as random-resized-crop and color-jitter, have proven critical to the success of contrastive learning techni…