activity
20242026
most citedERNIE 5.0 Technical Report

2 citations · 3 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2026

POINTS-GUI-G: GUI-Grounding Journey

Zhongyin Zhao, Yuan Liu, Yikun Liu +7

The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight…

cs.CL20262 cited

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CV2025

Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification

Yao Qin, Yangyang Yan, YuanChao Yang +4

Deep learning models have achieved remarkable success in medical image analysis but are fundamentally constrained by the requirement for large-scale, meticulously annotated dataset…

cs.CV2025

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li +11

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…

cs.LG2025

BigBang-Proton Technical Report: Next-Word-Prediction is Scientific Multitask Learner

Hengkui Wu, Liujiang Liu, Jihua He +23

We introduce BigBang-Proton, a unified sequence-based architecture for auto-regressive language modeling pretrained on cross-scale, cross-structure, cross-discipline real-world sci…

cs.CV2025

POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion

Yuan Liu, Zhongyin Zhao, Le Tian +8

High-quality labeled data is essential for training accurate document conversion models, particularly in domains with complex formats such as tables, formulas, and multi-column tex…