82 citations · 85 across the 16 of their papers we have counts for
5 papers · 1 filter
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
Junyang Wang, Haiyang Xu, Xi Zhang +4
Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from a fundamental conflict betwe…
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
Yusheng Song, Lirong Qiu, Xi Zhang +1
The detection of sophisticated hallucinations in Large Language Models (LLMs) is hampered by a ``Detection Dilemma'': methods probing internal states (Internal State Probing) excel…
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
Junyang Wang, Haiyang Xu, Xi Zhang +4
The exponential rise in mobile device usage necessitates streamlined automation for effective task management, yet many AI frameworks fall short due to inadequate operational exper…
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
Zhenhailong Wang, Haiyang Xu, Junyang Wang +5
Smartphones have become indispensable in modern life, yet navigating complex tasks on mobile devices often remains frustrating. Recent advancements in large multimodal model (LMM)-…
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
Junyang Wang, Haiyang Xu, Haitao Jia +6
Mobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario. Current Multi-modal Large Language Models (MLLMs), constrained by their train…