3 papers
cs.CL2026
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif +2
Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that…
cs.CL2026
The widening evaluation gap in medical large language model research 2023 to 2026
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed retu…
physics.soc-ph2026
Benchmarking large language model agent societies against human behavioural distributions
Raad Bin Tareaf
Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they st…