activity
20212024
most citedMERLOT: Multimodal Neural Script Knowledge Models

54 citations · 76 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2023

Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Wanrong Zhu, Jack Hessel, Anas Awadalla +7

In-context vision and language models like Flamingo support arbitrarily interleaved sequences of images and text as input. This format not only enables few-shot learning via interl…

cs.CV20222 cited

Learning Joint Representation of Human Motion and Language

Jihoon Kim, Youngjae Yu, Seungyoun Shin +2

In this work, we present MoLang (a Motion-Language connecting model) for learning joint representation of human motion and language, leveraging both unpaired and paired datasets of…

cs.CV2021

Pano-AVQA: Grounded Audio-Visual Question Answering on 360 Videos

Heeseung Yun, Youngjae Yu, Wonsuk Yang +2

360 videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial…

cs.CV20212 cited

Cycled Compositional Learning between Images and Text

Jongseok Kim, Youngjae Yu, Seunghwan Lee +1

We present an approach named the Cycled Composition Network that can measure the semantic distance of the composition of image-text embedding. First, the Composition Network transi…

cs.CV202154 cited

MERLOT: Multimodal Neural Script Knowledge Models

Rowan Zellers, Ximing Lu, Jack Hessel +5

As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce M…

cs.CV2021

ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning

Sangho Lee, Jiwan Chung, Youngjae Yu +4

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the eve…