3 papers
cs.AI2026
QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence
Zhichao Lin, Zhichao Liang, Gaoqiang Liu +6
As agentic foundation models continue to evolve, how to further improve their performance in vertical domains has become an important challenge. To this end, building upon Tongyi D…
cs.AI2025
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
Yunhe Yan, Shihe Wang, Jiajun Du +12
(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GU…
cs.SE2025
Every Software as an Agent: Blueprint and Case Study
Mengwei Xu
The rise of (multimodal) large language models (LLMs) has shed light on software agent -- where software can understand and follow user instructions in natural language. However, e…