19 citations · 19 across the 3 of their papers we have counts for
5 papers · 1 filter
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Xingyu Fu, Minqian Liu, Zhengyuan Yang +6
Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning s…
TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6
In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answ…
RePr: Improved Training of Convolutional Filters
Aaditya Prakash, James Storer, Dinei Florencio +1
A well-trained Convolutional Neural Network can easily be pruned without significant loss of performance. This is because of unnecessary overlap in the features captured by the net…
A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet Domain
Shuai Li, Dinei Florencio, Wanqing Li +2
Detecting camouflaged moving foreground objects has been known to be difficult due to the similarity between the foreground objects and the background. Conventional methods cannot…
Foreground Detection in Camouflaged Scenes
Shuai Li, Dinei Florencio, Yaqin Zhao +2
Foreground detection has been widely studied for decades due to its importance in many practical applications. Most of the existing methods assume foreground and background show vi…