3 papers
cs.AI2026
Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence
Yanan Wang, Shuaicong Hu, Jian Liu +3
The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-specific medical specialist model…
cs.AI2026
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
Yunqi Liu, Tong Niu, Zitong Wang +4
As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic inte…
stat.AP2025
Bridging the Data Gap in AI Reliability Research and Establishing DR-AIR, a Comprehensive Data Repository for AI Reliability
Simin Zheng, Jared M. Clark, Fatemeh Salboukh +10
Artificial intelligence (AI) technology and systems have been advancing rapidly. However, ensuring the reliability of these systems is crucial for fostering public confidence in th…