48 citations · 145 across the 12 of their papers we have counts for
7 papers · 1 filter
WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset
Andrea Burns, Krishna Srinivasan, Joshua Ainslie +5
Webpages have been a rich resource for language and vision-language tasks. Yet only pieces of webpages are kept: image-caption pairs, long text articles, or raw HTML, never all in…
A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding
Andrea Burns, Krishna Srinivasan, Joshua Ainslie +5
Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text…
Knowledge Prompts: Injecting World Knowledge into Language Models through Soft Prompts
Cicero Nogueira dos Santos, Zhe Dong, Daniel Cer +4
Soft prompts have been recently proposed as a tool for adapting large frozen language models (LMs) to new tasks. In this work, we repurpose soft prompts to the task of injecting wo…
Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma +7
Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other tasks where supervision is lim…
Neural Passage Retrieval with Improved Negative Contrast
Jing Lu, Gustavo Hernandez Abrego, Ji Ma +2
In this paper we explore the effects of negative sampling in dual encoder models used to retrieve passages for automatic question answering. We explore four negative sampling strat…
Interview: A Large-Scale Open-Source Corpus of Media Dialog
Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni +1
Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K convers…