activity
20242026
most citedLDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining

1 citations · 2 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CV2026

StrucTab: A Structured Optimization Framework for Table Parsing

Gengluo Li, Shangpin Peng, Chengquan Zhang +10

Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial layouts and textual content.…

cs.CV2026

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

Gengluo Li, Pengyuan Lyu, Chengquan Zhang +7

Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend…

cs.CV2026

MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

Gengluo Li, Chengquan Zhang, Yupu Liang +9

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding.…

cs.CV2025

HunyuanOCR Technical Report

Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +29

This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Na…

cs.CV2025

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu +5

Video text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field…

cs.CV2025

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

Gengluo Li, Huawen Shen, Yu Zhou

Chinese scene text retrieval is a practical task that aims to search for images containing visual instances of a Chinese query text. This task is extremely challenging because Chin…