activity
20242026
most citedERNIE 5.0 Technical Report

2 citations · 3 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.PF2026

How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding

Minghua He, Lingzhe Zhang, Yuan Liu +2

Parallel decoding improves generation efficiency by processing multiple decode positions within a single decode forward, but reported speedups conflate algorithmic token utilizatio…

cs.CV2026

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

Linhao Zhang, Aiwei Liu, Yuan Liu +1

Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences remains limited. We study this p…

cs.CV2026

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management

Yikun Liu, Yuan Liu, Haicheng Wang +6

Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…

cs.CV2026

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

Yikun Liu, Yuan Liu, Shangzhe Di +8

Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…

cs.CV2026

POINTS-GUI-G: GUI-Grounding Journey

Zhongyin Zhao, Yuan Liu, Yikun Liu +7

The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight…

cs.CL2026★ 2 cited

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…