3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 1 cited
Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA
Yongxin Zhu, Zhen Liu, Yukang Liang +4
In this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering.…
cs.CV2023★ 3 cited
Disorder-invariant Implicit Neural Representation
Hao Zhu, Shaowen Xie, Zhen Liu +6
Implicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problem…