9 papers
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States
Chen Liu, Ling Chen, Hanzhang Zhou +7
Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Existing methods treat memory l…
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding
Chen Liu, Ling Chen, Hanzhang Zhou +5
MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong instruction-following and se…
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity
Tianchen Guo, Chen Liu, Ling Chen +1
Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view human-centric scenes remain…
OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios
Xinyi Li, Zhen Fang, Yongxin Deng +12
Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challenges: inconsistent inference con…
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
Yuanwei Hu, Bo Peng, Yadan Luo +3
Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected inputs from unknown classes.…
Delving into Spectral Clustering with Vision-Language Representations
Bo Peng, Yuanwei Hu, Bo Liu +3
Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving…