6 citations · 6 across the 4 of their papers we have counts for
4 papers
Unhackable Temporal Rewarding for Scalable Video MLLMs
En Yu, Kangheng Lin, Liang Zhao +8
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. Th…
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
Zining Zhu, Liang Zhao, Kangheng Lin +7
This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained mul…
Slow Perception: Let's Perceive Geometric Figures Step-by-step
Haoran Wei, Youyang Yin, Yumeng Li +6
Recently, "visual o1" began to enter people's vision, with expectations that this slow-thinking design can solve visual reasoning tasks, especially geometric math problems. However…
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Haoran Wei, Chenglong Liu, Jinyue Chen +9
Traditional OCR systems (OCR-1.0) are increasingly unable to meet people's usage due to the growing demand for intelligent processing of man-made optical characters. In this paper,…