Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
cs.AI2025
Learning to Be A Doctor: Searching for Effective Medical Agent Architectures
Yangyang Zhuang, Wenjia Jiang, Jiayu Zhang +3
Large Language Model (LLM)-based agents have demonstrated strong capabilities across a wide range of tasks, and their application in the medical domain holds particular promise due…
cs.AI2025
AppAgentX: Evolving GUI Agents as Proficient Smartphone Users
Wenjia Jiang, Yangyang Zhuang, Chenxi Song +3
Recent advancements in Large Language Models (LLMs) have led to the development of intelligent LLM-based agents capable of interacting with graphical user interfaces (GUIs). These…