3 papers
cs.AI2026
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
Yunfei Feng, Xi Zhao, Cheng Zhang +5
Mobile agents can autonomously complete user-assigned tasks through GUI interactions. However, existing mainstream evaluation benchmarks, such as AndroidWorld, operate by connectin…
cs.CL2026
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
Zhaolu Kang, Junhao Gong, Wenqing Hu +15
Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowl…
cs.AI2025
Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
Zibin Liu, Cheng Zhang, Xi Zhao +6
Large Language Model (LLM) agents are increasingly deployed to automate complex workflows in mobile and desktop environments. However, current model-centric agent architectures str…