activity
20182022
most citedViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents

3 citations · 3 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2022

APRNet: Attention-based Pixel-wise Rendering Network for Photo-Realistic Text Image Generation

Yangming Shi, Haisong Ding, Kai Chen +1

Style-guided text image generation tries to synthesize text image by imitating reference image's appearance while keeping text content unaltered. The text image appearance includes…

cs.CL20213 cited

ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents

Weihong Lin, Qifang Gao, Lei Sun +4

Recent grid-based document representations like BERTgrid allow the simultaneous encoding of the textual and layout information of a document in a 2D feature map so that state-of-th…

cs.CL2020

A Study on Effects of Implicit and Explicit Language Model Information for DBLSTM-CTC Based Handwriting Recognition

Qi Liu, Lijuan Wang, Qiang Huo

Deep Bidirectional Long Short-Term Memory (D-BLSTM) with a Connectionist Temporal Classification (CTC) output layer has been established as one of the state-of-the-art solutions fo…

cs.CV2020

ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks

Chixiang Ma, Lei Sun, Zhuoyao Zhong +1

We introduce a new arbitrary-shaped text detection approach named ReLaText by formulating text detection as a visual relationship detection problem. To demonstrate the effectivenes…

cs.CV2018

Mask R-CNN with Pyramid Attention Network for Scene Text Detection

Zhida Huang, Zhuoyao Zhong, Lei Sun +1

In this paper, we present a new Mask R-CNN based text detection approach which can robustly detect multi-oriented and curved text from natural scene images in a unified manner. To…

cs.CV2018

An Anchor-Free Region Proposal Network for Faster R-CNN based Text Detection Approaches

Zhuoyao Zhong, Lei Sun, Qiang Huo

The anchor mechanism of Faster R-CNN and SSD framework is considered not effective enough to scene text detection, which can be attributed to its IoU based matching criterion betwe…