activity
20232026
most citedGeneral OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

6 citations · 9 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

STEP3-VL-10B Technical Report

Ailin Huang, Chengyuan Yao, Chunrui Han +90

We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…

cs.CV2025

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

Yana Wei, Liang Zhao, Jianjian Sun +15

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates…

cs.CV2025

DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning

Dongxu Liu, Jiahui Zhu, Yuang Peng +6

Autoencoders empower state-of-the-art image and video generative models by compressing pixels into a latent space through visual tokenization. Although recent advances have allevia…

cs.CV2025

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

En Yu, Kangheng Lin, Liang Zhao +11

Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, ou…

cs.CV2025

Step1X-Edit: A Practical Framework for General Image Editing

Shiyu Liu, Yucheng Han, Peng Xing +21

In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has i…

cs.CV20246 cited

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Haoran Wei, Chenglong Liu, Jinyue Chen +9

Traditional OCR systems (OCR-1.0) are increasingly unable to meet people's usage due to the growing demand for intelligent processing of man-made optical characters. In this paper,…