8 papers
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
Yubin Kim, Chanwoo Park, Taehan Kim +9
Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls. Without enforcement, a team…
Uneven Evolution of Cognition Across Generations of Generative AI Models
Isaac Galatzer-Levy, Daniel McDuff, Xin Liu +1
The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task performance. Here, we introduce a…
How Well Do Multimodal Models Reason on ECG Signals?
Maxwell A. Xu, Harish Haresamudram, Catherine W. Liu +11
While multimodal large language models offer a promising solution to the "black box" nature of health AI by generating interpretable reasoning traces, verifying the validity of the…
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
Yubin Kim, Hyewon Jeong, Shan Chen +24
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence an…
Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety
Yubin Kim, Hyewon Jeong, Chanwoo Park +9
Large language models (LLMs) deployed as agents introduce significant safety risks in clinical settings due to their potential for error and single points of failure. We introduce…
VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation
Yubin Kim, Taehan Kim, Wonjune Kang +8
Vocal health plays a crucial role in peoples' lives, significantly impacting their communicative abilities and interactions. However, despite the global prevalence of voice disorde…