activity
20142023
most citedTextBoxes: A Fast Text Detector with a Single Deep Neural Network

445 citations · 780 across the 29 of their papers we have counts for

collaborators

42 papers

cs.CV2024

Bridging the Gap Between End-to-End and Two-Step Text Spotting

Mingxin Huang, Hongliang Li, Yuliang Liu +2

Modularity plays a crucial role in the development and maintenance of complex systems. While end-to-end text spotting efficiently mitigates the issues of error accumulation and sub…

cs.CV20241 cited

Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis

Xin Zhou, Dingkang Liang, Wei Xu +4

Point cloud analysis has achieved outstanding performance by transferring point cloud pre-trained models. However, existing methods for model adaptation usually update all model pa…

cs.CV20241 cited

OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Jianqiang Wan, Sibo Song, Wenwen Yu +6

Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Gene…

cs.CV2024

PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Zheng Zhang, Yeyao Ma, Enming Zhang +1

PSALM is a powerful extension of the Large Multi-modal Model (LMM) to address the segmentation task challenges. To overcome the limitation of the LMM being limited to textual outpu…

cs.CV202412 cited

TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Yuliang Liu, Biao Yang, Qiang Liu +4

We present TextMonkey, a large multimodal model (LMM) tailored for text-centric tasks. Our approach introduces enhancement across several dimensions: By adopting Shifted Window Att…

cs.CV20242 cited

Anomaly Detection by Adapting a pre-trained Vision Language Model

Yuxuan Cai, Xinwei He, Dingkang Liang +2

Recently, large vision and language models have shown their success when adapting them to many downstream tasks. In this paper, we present a unified framework named CLIP-ADA for An…