collaborators

6 papers

cs.CV2026

H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification

Yongji Zhang, Siqi Li, Kuiyang Huang +2

Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on f…

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CV2025

PaddleOCR 3.0 Technical Report

Cheng Cui, Ting Sun, Manhui Lin +16

This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the…

cs.CV2025

PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks

Feng Ni, Kui Huang, Yao Lu +4

With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for…

cs.CV2025

PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding

Kui Huang, Xinrong Chen, Wenyu Lv +3

This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-D…

cs.CV2025

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information

Xu Chu, Xinrong Chen, Guanyu Wang +5

Inference time scaling drives extended reasoning to enhance the performance of Vision-Language Models (VLMs), thus forming powerful Vision-Language Reasoning Models (VLRMs). Howeve…