445 citations · 771 across the 24 of their papers we have counts for
15 papers · 1 filter
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
Zhang Li, Yuliang Liu, Qiang Liu +8
We introduce MonkeyOCR, a document parsing model that advances the state of the art by leveraging a Structure-Recognition-Relation (SRR) triplet paradigm. This design simplifies wh…
Side Adapter Network for Open-Vocabulary Semantic Segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei +2
This paper presents a new framework for open-vocabulary semantic segmentation with the pre-trained vision-language model, named Side Adapter Network (SAN). Our approach models the…
CAPE: Camera View Position Embedding for Multi-View 3D Object Detection
Kaixin Xiong, Shi Gong, Xiaoqing Ye +5
In this paper, we address the problem of detecting 3D objects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the geometric…
Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition
Xudong Xie, Ling Fu, Zhifei Zhang +2
Artistic text recognition is an extremely challenging task with a wide range of applications. However, current scene text recognition methods mainly focus on irregular text while h…
Optimal Boxes: Boosting End-to-End Scene Text Recognition by Adjusting Annotated Bounding Boxes via Reinforcement Learning
Jingqun Tang, Wenming Qian, Luchuan Song +3
Text detection and recognition are essential components of a modern OCR system. Most OCR approaches attempt to obtain accurate bounding boxes of text at the detection stage, which…
When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition
Bohan Li, Ye Yuan, Dingkang Liang +5
Recently, most handwritten mathematical expression recognition (HMER) methods adopt the encoder-decoder networks, which directly predict the markup sequences from formula images wi…