489 citations · 500 across the 3 of their papers we have counts for
3 papers
Making the Most of What You Have: Adapting Pre-trained Visual Language Models in the Low-data Regime
Chuhan Zhang, Antoine Miech, Jiajun Shen +2
Large-scale visual language models are widely used as pre-trained models and then adapted for various downstream tasks. While humans are known to efficiently learn new tasks from a…
Zorro: the masked multimodal transformer
Adrià Recasens, Jason Lin, Joāo Carreira +8
Attention-based models are appealing for multimodal processing because inputs from multiple modalities can be concatenated and fed to a single backbone network - thus requiring ver…
Semantic Segmentation using Adversarial Networks
Pauline Luc, Camille Couprie, Soumith Chintala +1
Adversarial training has been shown to produce state of the art results for generative image modeling. In this paper we propose an adversarial training approach to train semantic s…