activity
20222024
most citedDatasets for Large Language Models: A Comprehensive Survey

16 citations · 33 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2024

Predicting the Original Appearance of Damaged Historical Documents

Zhenhua Yang, Dezhi Peng, Yongxin Shi +3

Historical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing…

cs.CL202416 cited

Datasets for Large Language Models: A Comprehensive Survey

Yang Liu, Jiahuan Cao, Chongyu Liu +2

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the found…

cs.CV202312 cited

Exploring OCR Capabilities of GPT-4V(ision) : A Quantitative and In-depth Evaluation

Yongxin Shi, Dezhi Peng, Wenhui Liao +5

This paper presents a comprehensive evaluation of the Optical Character Recognition (OCR) capabilities of the recently released GPT-4V(ision), a Large Multimodal Model (LMM). We as…

cs.CV20234 cited

Revisiting Scene Text Recognition: A Data Perspective

Qing Jiang, Jiapeng Wang, Dezhi Peng +2

This paper aims to re-assess scene text recognition (STR) from a data-oriented perspective. We begin by revisiting the six commonly used benchmarks in STR and observe a trend of pe…

cs.CV20221 cited

Don't Forget Me: Accurate Background Recovery for Text Removal via Modeling Local-Global Context

Chongyu Liu, Lianwen Jin, Yuliang Liu +4

Text removal has attracted increasingly attention due to its various applications on privacy protection, document restoration, and text editing. It has shown significant progress w…