4 papers · 1 filter
Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA
Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto
Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. Geometrically, that implies that tru…
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer +32
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant…
Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study
Rachel M. Murphy, Nishant Mishra, Nicolette F. de Keizer +5
In this study, we establish a benchmark for adverse drug event (ADE) detection in Dutch clinical free-text documents using several transformer models, clinical scenarios, and fit-f…
LLM aided semi-supervision for Extractive Dialog Summarization
Nishant Mishra, Gaurav Sahu, Iacer Calixto +2
Generating high-quality summaries for chat dialogs often requires large labeled datasets. We propose a method to efficiently use unlabeled data for extractive summarization of cust…