activity
20212023
most citedPaLI-X: On Scaling up a Multilingual Vision and Language Model

39 citations · 101 across the 14 of their papers we have counts for

collaborators

14 papers

cs.CV20231 cited

AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description

Tengda Han, Max Bain, Arsha Nagrani +3

Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presen…

cs.CV20232 cited

VidChapters-7M: Video Chapters at Scale

Antoine Yang, Arsha Nagrani, Ivan Laptev +2

Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly…

cs.CL202311 cited

LanSER: Language-Model Supported Speech Emotion Recognition

Taesik Gong, Josh Belanich, Krishna Somandepalli +3

Speech emotion recognition (SER) models typically rely on costly human-labeled data for training, making scaling methods to large speech datasets and nuanced emotion taxonomies dif…

cs.CV20232 cited

UnLoc: A Unified Framework for Video Localization Tasks

Shen Yan, Xuehan Xiong, Arsha Nagrani +5

While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos…

cs.CL2023

Modular Visual Question Answering via Code Generation

Sanjay Subramanian, Medhini Narasimhan, Kushal Khangaonkar +6

We present a framework that formulates visual question answering as modular code generation. In contrast to prior work on modular approaches to VQA, our approach requires no additi…

cs.CV202339 cited

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…