72 citations · 346 across the 19 of their papers we have counts for
Showing 2022 · cs.CLShow all
2 papers · 2 filters
cs.CL2022★ 9 cited
MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
Wenhu Chen, Hexiang Hu, Xi Chen +2
While language Models store a massive amount of world knowledge implicitly in their parameters, even very large models often fail to encode information about rare entities and even…
cs.CL2022★ 47 cited
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Kenton Lee, Mandar Joshi, Iulia Turc +7
Visually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to th…