9 citations · 9 across the 2 of their papers we have counts for
2 papers
cs.CV2023
What Large Language Models Bring to Text-rich VQA?
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…
cs.AI2023★ 9 cited
UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding
Hao Feng, Zijian Wang, Jingqun Tang +4
In the era of Large Language Models (LLMs), tremendous strides have been made in the field of multimodal understanding. However, existing advanced algorithms are limited to effecti…