activity
20232026
most citedTextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

12 citations · 17 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Yiran Guan, Sifan Tu, Dingkang Liang +6

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel…

cs.CV2026

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

Hanshen Zhu, Yuliang Liu, Xuecheng Wu +7

Visual Text Rendering (VTR) remains a critical challenge in text-to-image generation, where even advanced models frequently produce text with structural anomalies such as distortio…

cs.CV2025

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Ling Fu, Zhebin Kuang, Jiajun Song +21

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive p…

cs.CV2024

MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark

Bin Shan, Xiang Fei, Wei Shi +6

The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchma…

cs.CV2024

Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering

Hiba Maryam, Ling Fu, Jiajun Song +4

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and lingu…

cs.CV2024

The First Swahili Language Scene Text Detection and Recognition Dataset

Fadila Wendigoundi Douamba, Jianjun Song, Ling Fu +2

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals wi…