2 citations · 3 across the 13 of their papers we have counts for
6 papers · 1 filter
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
Minghua He, Lingzhe Zhang, Yuan Liu +2
Parallel decoding improves generation efficiency by processing multiple decode positions within a single decode forward, but reported speedups conflate algorithmic token utilizatio…
DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?
Linhao Zhang, Aiwei Liu, Yuan Liu +1
Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences remains limited. We study this p…
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Yikun Liu, Yuan Liu, Haicheng Wang +6
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Yikun Liu, Yuan Liu, Shangzhe Di +8
Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…
POINTS-GUI-G: GUI-Grounding Journey
Zhongyin Zhao, Yuan Liu, Yikun Liu +7
The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…