activity
20182022
most citedLearning Audio-Video Modalities from Image Captions

6 citations · 6 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CV2022

AVATAR submission to the Ego4D AV Transcription Challenge

Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

In this report, we describe our submission to the Ego4D AudioVisual (AV) Speech Transcription Challenge 2022. Our pipeline is based on AVATAR, a state of the art encoder-decoder mo…

cs.CV20226 cited

Learning Audio-Video Modalities from Image Captions

Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold +4

A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of s…

cs.CV2020

Look Before you Speak: Visually Contextualized Utterances

Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

While most conversational AI systems focus on textual dialogue only, conditioning utterances on visual context (when it's available) can lead to more realistic conversations. Unfor…

cs.CV2019

Reinforcing an Image Caption Generator Using Off-Line Human Feedback

Paul Hongsuck Seo, Piyush Sharma, Tomer Levinboim +2

Human ratings are currently the most accurate way to assess the quality of an image captioning model, yet most often the only used outcome of an expensive human rating evaluation i…

cs.LG2019

Regularizing Neural Networks via Stochastic Branch Layers

Wonpyo Park, Paul Hongsuck Seo, Bohyung Han +1

We introduce a novel stochastic regularization technique for deep neural networks, which decomposes a layer into multiple branches with different parameters and merges stochastical…

cs.LG2018

Learning for Single-Shot Confidence Calibration in Deep Neural Networks through Stochastic Inferences

Seonguk Seo, Paul Hongsuck Seo, Bohyung Han

We propose a generic framework to calibrate accuracy and confidence of a prediction in deep neural networks through stochastic inferences. We interpret stochastic regularization us…