12 papers
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
Minghua He, Lingzhe Zhang, Yuan Liu +2
Parallel decoding improves generation efficiency by processing multiple decode positions within a single decode forward, but reported speedups conflate algorithmic token utilizatio…
DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?
Linhao Zhang, Aiwei Liu, Yuan Liu +1
Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences remains limited. We study this p…
Enhancing Speech Large Language Models through Reinforced Behavior Alignment
Yansong Liu, Jiateng Li, Yuan Liu
The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyond text to other modalities, whic…
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Yikun Liu, Yuan Liu, Le Tian +6
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Yikun Liu, Yuan Liu, Shangzhe Di +8
Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…
POINTS-GUI-G: GUI-Grounding Journey
Zhongyin Zhao, Yuan Liu, Yikun Liu +7
The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight…