4 papers
Does Chain-of-Thought Reasoning Help Mobile GUI Agent? An Empirical Study
Li Zhang, Longxi Gao, Mengwei Xu
Reasoning capabilities have significantly improved the performance of vision-language models (VLMs) in domains such as mathematical problem-solving, coding, and visual question-ans…
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
Haiyang Shen, Yue Li, Desong Meng +5
Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recen…
DroidCall: A Dataset for LLM-powered Android Intent Invocation
Weikai Xie, Li Zhang, Shihe Wang +2
The growing capabilities of large language models in natural language understanding significantly strengthen existing agentic systems. To power performant on-device mobile agents f…
A First Look at GPT Apps: Landscape and Vulnerability
Zejun Zhang, Li Zhang, Xin Yuan +3
Following OpenAI's introduction of GPTs, a surge in GPT apps has led to the launch of dedicated LLM app stores. Nevertheless, given its debut, there is a lack of sufficient underst…