4 papers
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
Xiaoyi Chen, Yifei Gao, Yang Xu +3
Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defect…
TrialCalibre: A Fully Automated Causal Engine for RCT Benchmarking and Observational Trial Calibration
Amir Habibdoust, Xing Song
Real-world evidence (RWE) studies that emulate target trials increasingly inform regulatory and clinical decisions, yet residual, hard-to-quantify biases still limit their credibil…
AgenticAD: A Specialized Multiagent System Framework for Holistic Alzheimer Disease Management
Adib Bazgir, Amir Habibdoust, Xing Song +1
Alzheimer's disease (AD) presents a complex, multifaceted challenge to patients, caregivers, and the healthcare system, necessitating integrated and dynamic support solutions. Whil…
Causal MAS: A Survey of Large Language Model Architectures for Discovery and Effect Estimation
Adib Bazgir, Amir Habibdoust, Yuwen Zhang +1
Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning and generation tasks. However, their proficiency in complex causal reasoning, discovery,…