activity
20182024
most citedScaling Up Models and Data with and

48 citations · 145 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20231 cited

WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset

Andrea Burns, Krishna Srinivasan, Joshua Ainslie +5

Webpages have been a rich resource for language and vision-language tasks. Yet only pieces of webpages are kept: image-caption pairs, long text articles, or raw HTML, never all in…

cs.CL2023

A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding

Andrea Burns, Krishna Srinivasan, Joshua Ainslie +5

Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text…

cs.CL20223 cited

Knowledge Prompts: Injecting World Knowledge into Language Models through Soft Prompts

Cicero Nogueira dos Santos, Zhe Dong, Daniel Cer +4

Soft prompts have been recently proposed as a tool for adapting large frozen language models (LMs) to new tasks. In this work, we repurpose soft prompts to the task of injecting wo…

cs.CL202247 cited

Promptagator: Few-shot Dense Retrieval From 8 Examples

Zhuyun Dai, Vincent Y. Zhao, Ji Ma +7

Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other tasks where supervision is lim…

cs.CL2020

Neural Passage Retrieval with Improved Negative Contrast

Jing Lu, Gustavo Hernandez Abrego, Ji Ma +2

In this paper we explore the effects of negative sampling in dual encoder models used to retrieve passages for automatic question answering. We explore four negative sampling strat…

cs.CL20205 cited

Interview: A Large-Scale Open-Source Corpus of Media Dialog

Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni +1

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K convers…