activity
20242026
most citedAppAgent: Multimodal Agents as Smartphone Users

11 citations · 11 across the 1 of their papers we have counts for

collaborators

9 papers

cs.CV202611 cited

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +6

Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper introduces a novel LLM-based mult…

cs.SD2026

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models

Yanda Li, Yuhan Liu, Zirui Song +3

Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be un…

cs.HC2025

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

Yanda Li, Chi Zhang, Wenjia Jiang +6

With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user int…

cs.AI2025

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang +3

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and pro…

cs.CL2025

SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

Wanqi Yang, Yanda Li, Yunchao Wei +2

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surf…

cs.SD2025

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models

Wanqi Yang, Yanda Li, Meng Fang +2

Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research foc…