activity
20172021
most citedMERLOT: Multimodal Neural Script Knowledge Models

54 citations · 60 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV202154 cited

MERLOT: Multimodal Neural Script Knowledge Models

Rowan Zellers, Ximing Lu, Jack Hessel +5

As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce M…

cs.CV2020

Identity-Aware Multi-Sentence Video Description

Jae Sung Park, Trevor Darrell, Anna Rohrbach

Standard video and movie description tasks abstract away from person identities, thus failing to link identities across sentences. We propose a multi-sentence Identity-Aware Video…

cs.CV2020

VisualCOMET: Reasoning about the Dynamic Context of a Still Image

Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi +2

Even from a single frame of a still image, people can reason about the dynamic story of the image before, after, and beyond the frame. For example, given an image of a man struggli…

cs.CV2018

Adversarial Inference for Multi-Sentence Video Description

Jae Sung Park, Marcus Rohrbach, Trevor Darrell +1

While significant progress has been made in the image captioning task, video description is still in its infancy due to the complex nature of video data. Generating multi-sentence…

cs.CV20172 cited

Generation of High Dynamic Range Illumination from a Single Image for the Enhancement of Undesirably Illuminated Images

Jae Sung Park, Nam Ik Cho

This paper presents an algorithm that enhances undesirably illuminated images by generating and fusing multi-level illuminations from a single image.The input image is first decomp…