2 citations · 4 across the 10 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
Yi Zhu, Xiongwei Wu, Qiyi Wang +8
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet…
cs.AI2026
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
Ming Ma, Yi Zhu, Yiran Zhong +6
Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing…