activity
20182025
most citedTAP: Text-Aware Pre-training for Text-VQA and Text-Caption

19 citations · 19 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Xingyu Fu, Minqian Liu, Zhengyuan Yang +6

Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning s…

cs.CV2023

Diffusion-based Document Layout Generation

Liu He, Yijuan Lu, John Corring +2

We develop a diffusion-based approach for various document layout sequence generation. Layout sequences specify the contents of a document design in an explicit format. Our novel d…

cs.CV202019 cited

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answ…

cs.CV2020

Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language

Songyang Zhang, Houwen Peng, Jianlong Fu +2

We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the contex…

cs.CV2018

FANet: Quality-Aware Feature Aggregation Network for Robust RGB-T Tracking

Yabin Zhu, Chenglong Li, Bin Luo +1

This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose…

cs.CV2018

RGB-T Object Tracking:Benchmark and Baseline

Chenglong Li, Xinyan Liang, Yijuan Lu +2

RGB-Thermal (RGB-T) object tracking receives more and more attention due to the strongly complementary benefits of thermal information to visible data. However, RGB-T research is l…