activity
20242026
most citedUPOCR: Towards Unified Pixel-Level OCR Interface

1 citations · 2 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CV20261 cited

UPOCR: Towards Unified Pixel-Level OCR Interface

Dezhi Peng, Zhenhua Yang, Jiaxin Zhang +5

Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases t…

cs.CV20261 cited

OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities

Peirong Zhang, Haowei Xu, Jiaxin Zhang +7

Improving visual text synthesis has long been a challenging and evolving frontier for image generation models. While recent state-of-the-art (SOTA) models have made remarkable stri…

cs.CV2025

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Hongliang Li, Jiaxin Zhang, Wenhui Liao +3

Current Multimodal Large Language Model (MLLM) architectures face a critical tradeoff between performance and efficiency: decoder-only architectures achieve higher performance but…

cs.LG2025

Smaller But Better: Unifying Layout Generation with Smaller Large Language Models

Peirong Zhang, Jiaxin Zhang, Jiahuan Cao +2

We propose LGGPT, an LLM-based model tailored for unified layout generation. First, we propose Arbitrary Layout Instruction (ALI) and Universal Layout Response (ULR) as the uniform…

cs.CV2024

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Jiaxin Zhang, Wentao Yang, Songxuan Lai +2

Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layout…

cs.CV2024

LEGO: Self-Supervised Representation Learning for Scene Text Images

Yujin Ren, Jiaxin Zhang, Lianwen Jin

In recent years, significant progress has been made in scene text recognition by data-driven methods. However, due to the scarcity of annotated real-world data, the training of the…