297 citations · 1k across the 40 of their papers we have counts for
8 papers · 1 filter
Connecting Vision and Language with Localized Narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo +2
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneo…
Reinforcing an Image Caption Generator Using Off-Line Human Feedback
Paul Hongsuck Seo, Piyush Sharma, Tomer Levinboim +2
Human ratings are currently the most accurate way to assess the quality of an image captioning model, yet most often the only used outcome of an expensive human rating evaluation i…
A Case Study on Combining ASR and Visual Features for Generating Instructional Video Captions
Jack Hessel, Bo Pang, Zhenhai Zhu +1
Instructional videos get high-traffic on video sharing platforms, and prior work suggests that providing time-stamped, subtask annotations (e.g., "heat the oil in the pan") improve…
Multi-stage Pretraining for Abstractive Summarization
Sebastian Goodman, Zhenzhong Lan, Radu Soricut
Neural models for abstractive summarization tend to achieve the best performance in the presence of highly specialized, summarization specific modeling add-ons such as pointer-gene…
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
Soravit Changpinyo, Bo Pang, Piyush Sharma +1
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster…
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman +3
Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases be…