6 papers
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Jiancheng Huang, Gengwei Zhang, Zequn Jie +5
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal spa…
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Mengyu Zheng, Kai Han, Boxun Li +13
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not b…
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
Yanda Li, Chi Zhang, Wenjia Jiang +6
With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user int…
Foundations and Recent Trends in Multimodal Mobile Agents: A Survey
Biao Wu, Yanda Li, Zhiwei Zhang +3
Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and pro…
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
Wanqi Yang, Yanda Li, Yunchao Wei +2
Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surf…
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
Wanqi Yang, Yanda Li, Meng Fang +2
Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research foc…