4 citations · 6 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6
Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…
cs.CV2023★ 4 cited
SimCon Loss with Multiple Views for Text Supervised Semantic Segmentation
Yash Patel, Yusheng Xie, Yi Zhu +2
Learning to segment images purely by relying on the image-text alignment from web data can lead to sub-optimal performance due to noise in the data. The noise comes from the sample…
cs.CV2021★ 2 cited
LaTr: Layout-Aware Transformer for Scene-Text VQA
Ali Furkan Biten, Ron Litman, Yusheng Xie +2
We propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over…