2 citations · 2 across the 4 of their papers we have counts for
8 papers
TRACER: Early Failure Detection for Task-Oriented Dialogue
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present…
When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely viewed as a promising solutio…
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data
Amir Mousavi, Mohammad Sadegh Sirjani, Erfan Nourbakhsh +5
Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automat…
CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment
Amir Mousavi, Erfan Nourbakhsh, Mohammad Sadegh Sirjani +5
Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize p…
Prompting Underestimates LLM Capability for Time Series Classification
Dan Schumacher, Erfan Nourbakhsh, Rocky Slavin +1
Prompt-based evaluations suggest that large language models (LLMs) perform poorly on time series classification, raising doubts about whether they encode meaningful temporal struct…
An Analysis of Automated Use Case Component Extraction from Scenarios using ChatGPT
Pragyan KC, Rocky Slavin, Sepideh Ghanavati +2
Mobile applications (apps) are often developed by only a small number of developers with limited resources, especially in the early years of the app's development. In this setting,…