6 papers
Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes
Xiao-Hang Jiang, Han-Jie Guo, Ying-Si Liang +4
Ambient clinical scribes increasingly combine Automatic Speech Recognition with Large Language Models to automate documentation. However, traditional metrics like Word Error Rate m…
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
Lei Jiang, Yue Zhou, Natalie Parde
Reliable early detection of Alzheimer's disease (AD) is challenging, particularly due to the limited availability of labeled data. While large language models (LLMs) have shown str…
MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition
Yifan Bao, Xinyu Xi, Xinyu Liu +8
Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation p…
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
Chaodong Tong, Qi Zhang, Zhuojun Jiang +2
Large language models (LLMs) achieve strong question answering (QA) performance but can produce fluent answers unsupported by available evidence. Existing hallucination detectors o…
Knowledge Dependency Estimation for Reliable Question Answering
Chaodong Tong, Qi Zhang, Nannan Sun +2
Reliable question answering requires identifying not only whether an answer is correct, but also which available knowledge the prediction depends on. In realistic LLM-based QA, thi…
FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering
Chaodong Tong, Qi Zhang, Chen Li +2
Faithfulness hallucinations in VQA occur when vision-language models produce fluent yet visually ungrounded answers, severely undermining their reliability in safety-critical appli…