6 citations · 15 across the 6 of their papers we have counts for
4 papers · 1 filter
PBFormer: Capturing Complex Scene Text Shape with Polynomial Band Transformer
Ruijin Liu, Ning Lu, Dapeng Chen +3
We present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has…
ChartDETR: A Multi-shape Detection Network for Visual Chart Recognition
Wenyuan Xue, Dapeng Chen, Baosheng Yu +3
Visual chart recognition systems are gaining increasing attention due to the growing demand for automatically identifying table headers and values from chart images. Current method…
Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling
Yongshuai Huang, Ning Lu, Dapeng Chen +5
Table structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text appr…
Video Action Recognition with Attentive Semantic Units
Yifei Chen, Dapeng Chen, Ruijin Liu +2
Visual-Language Models (VLMs) have significantly advanced action video recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to le…