54 citations · 74 across the 4 of their papers we have counts for
5 papers
Multimodal Knowledge Alignment with Reinforcement Learning
Youngjae Yu, Jiwan Chung, Heeseung Yun +8
Large language models readily adapt to novel settings, even without task-specific training data. Can their zero-shot capacity be extended to multimodal inputs? In this work, we pro…
Pano-AVQA: Grounded Audio-Visual Question Answering on 360 Videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang +2
360 videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial…
Cycled Compositional Learning between Images and Text
Jongseok Kim, Youngjae Yu, Seunghwan Lee +1
We present an approach named the Cycled Composition Network that can measure the semantic distance of the composition of image-text embedding. First, the Composition Network transi…
MERLOT: Multimodal Neural Script Knowledge Models
Rowan Zellers, Ximing Lu, Jack Hessel +5
As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce M…
ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning
Sangho Lee, Jiwan Chung, Youngjae Yu +4
The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the eve…