◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Humen Zhong

10 papers hereh-index 66.3k citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author6

Across the 8 of 10 papers where every author was matched, so the position is known.

fields
  • cs.CV10

identity via Semantic Scholar / OpenAlex

activity
20212026
most citedQwen2.5-VL Technical Report

82 citations · 106 across the 9 of their papers we have counts for

collaborators
Showing 2024Show all

3 papers · 1 filter

cs.CV2024

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Zhibo Yang, Jun Tang, Zhaohai Li +9

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what exten…

cs.CV2024

VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer

Humen Zhong, Zhibo Yang, Zhaohai Li +4

Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towar…

cs.CV2024

Platypus: A Generalized Specialist Model for Reading Text in Various Forms

Peng Wang, Zhaohai Li, Jun Tang +4

Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. P…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.