1 paper · 1 filter
Jonathan Steinberg, Oren Gal
Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language processing stream? We investigate th…