31 citations · 49 across the 16 of their papers we have counts for
4 papers · 1 filter
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
Erfan Nourbakhsh, Ke Yang, Anthony Rios
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf…
TRACER: Early Failure Detection for Task-Oriented Dialogue
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present…
When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely viewed as a promising solutio…
Prompting Underestimates LLM Capability for Time Series Classification
Dan Schumacher, Erfan Nourbakhsh, Rocky Slavin +1
Prompt-based evaluations suggest that large language models (LLMs) perform poorly on time series classification, raising doubts about whether they encode meaningful temporal struct…