activity
20242026
collaborators

7 papers

cs.CL2026

Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces

Zijing Shi, Meng Fang, Ling Chen

As autonomous web agents are increasingly deployed to perform real-world tasks, ensuring their safety has become a critical concern. In this work, we study web agent behavior under…

cs.AI2025

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang +3

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and pro…

cs.CL2025

SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

Wanqi Yang, Yanda Li, Yunchao Wei +2

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surf…

cs.SD2025

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models

Wanqi Yang, Yanda Li, Meng Fang +2

Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research foc…

cs.AI2025

MMAC-Copilot: Multi-modal Agent Collaboration Operating Copilot

Zirui Song, Yaohang Li, Meng Fang +6

Large language model agents that interact with PC applications often face limitations due to their singular mode of interaction with real-world environments, leading to restricted…

cs.CL2025

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

Wanqi Yang, Yanda Li, Meng Fang +1

Understanding temporal dynamics is critical for conversational agents, enabling effective content analysis and informed decision-making. However, time-aware datasets, particularly…