most citedMCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents

2 citations · 3 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CL2025

Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution

Zhiyang Chen, Daliang Xu, Haiyang Shen +5

Performing Retrieval-Augmented Generation (RAG) directly on mobile devices is promising for data privacy and responsiveness but is hindered by the architectural constraints of mobi…

cs.AI20252 cited

MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents

Yunhe Yan, Shihe Wang, Jiajun Du +12

(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GU…

cs.LG2025

MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs

Zhenyan Lu, Daliang Xu, Dongqi Cai +5

Large language models (LLMs) are deployed on mobile devices to power killer applications such as intelligent assistants. LLMs pre-trained on general corpora often hallucinate when…

cs.LG2025

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades

Yanan Li, Fanxu Meng, Muhan Zhang +3

As Large Language Models (LLMs) are frequently updated, LoRA weights trained on earlier versions quickly become obsolete. The conventional practice of retraining LoRA weights from…

cs.IR2024

Recall: Empowering Multimodal Embedding for Edge Devices

Dongqi Cai, Shangguang Wang, Chen Peng +2

Human memory is inherently prone to forgetting. To address this, multimodal embedding models have been introduced, which transform diverse real-world data into a unified embedding…

cs.HC2024

MobileViews: A Million-scale and Diverse Mobile GUI Dataset

Longxi Gao, Li Zhang, Shihe Wang +6

Visual language models (VLMs) empower mobile GUI agents to interpret complex mobile screens and respond to user requests. Training such capable agents requires large-scale, high-qu…