most citedGeneral OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

6 citations · 7 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL20251 cited

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Ailin Huang, Boyong Wu, Bruce Wang +142

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such…

cs.CV2025

Unhackable Temporal Rewarding for Scalable Video MLLMs

En Yu, Kangheng Lin, Liang Zhao +8

In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. Th…

cs.AI2025

PerPO: Perceptual Preference Optimization via Discriminative Rewarding

Zining Zhu, Liang Zhao, Kangheng Lin +7

This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained mul…

cs.CV2025

Slow Perception: Let's Perceive Geometric Figures Step-by-step

Haoran Wei, Youyang Yin, Yumeng Li +6

Recently, "visual o1" began to enter people's vision, with expectations that this slow-thinking design can solve visual reasoning tasks, especially geometric math problems. However…

cs.CV2025

Taming Teacher Forcing for Masked Autoregressive Video Generation

Deyu Zhou, Quan Sun, Yuang Peng +8

We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation,…

cs.CV20246 cited

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Haoran Wei, Chenglong Liu, Jinyue Chen +9

Traditional OCR systems (OCR-1.0) are increasingly unable to meet people's usage due to the growing demand for intelligent processing of man-made optical characters. In this paper,…