3 citations · 5 across the 2 of their papers we have counts for
4 papers · 1 filter
Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
Yifei Chen, Dapeng Chen, Ruijin Liu +3
Large-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an "adapt then align" paradigm, which…
PBFormer: Capturing Complex Scene Text Shape with Polynomial Band Transformer
Ruijin Liu, Ning Lu, Dapeng Chen +3
We present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has…
ChartDETR: A Multi-shape Detection Network for Visual Chart Recognition
Wenyuan Xue, Dapeng Chen, Baosheng Yu +3
Visual chart recognition systems are gaining increasing attention due to the growing demand for automatically identifying table headers and values from chart images. Current method…
Self-distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach
Ziyin Zhang, Ning Lu, Minghui Liao +4
Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutive…