activity
20182021
most citedTAP: Text-Aware Pre-training for Text-VQA and Text-Caption

19 citations · 19 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL2021

LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Yiheng Xu, Tengchao Lv, Lei Cui +5

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually-rich document understanding tasks recently, which demonstrates the great potential f…

cs.CV202019 cited

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answ…

cs.CV2020

Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language

Songyang Zhang, Houwen Peng, Jianlong Fu +2

We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the contex…

cs.CV2018

FANet: Quality-Aware Feature Aggregation Network for Robust RGB-T Tracking

Yabin Zhu, Chenglong Li, Bin Luo +1

This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose…

cs.CV2018

RGB-T Object Tracking:Benchmark and Baseline

Chenglong Li, Xinyan Liang, Yijuan Lu +2

RGB-Thermal (RGB-T) object tracking receives more and more attention due to the strongly complementary benefits of thermal information to visible data. However, RGB-T research is l…