4 citations · 7 across the 2 of their papers we have counts for
6 papers · 1 filter
Evaluating the Zero-shot Robustness of Instruction-tuned Language Models
Jiuding Sun, Chantal Shaib, Byron C. Wallace
Instruction fine-tuning has recently emerged as a promising approach for improving the zero-shot capabilities of Large Language Models (LLMs) on new tasks. This technique has shown…
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung +4
Evaluating multi-document summarization (MDS) quality is difficult. This is especially true in the case of MDS for biomedical literature reviews, where models must synthesize contr…
Summarizing, Simplifying, and Synthesizing Medical Evidence Using GPT-3 (with Varying Success)
Chantal Shaib, Millicent L. Li, Sebastian Joseph +3
Large language models, particularly GPT-3, are able to produce high quality summaries of general domain news articles in few- and zero-shot settings. However, it is unclear if such…
Multilingual Simplification of Medical Texts
Sebastian Joseph, Kathryn Kazanas, Keziah Reina +4
Automated text simplification aims to produce simple versions of complex texts. This task is especially useful in the medical domain, where the latest medical findings are typicall…
Appraising the Potential Uses and Harms of LLMs for Medical Systematic Reviews
Hye Sun Yun, Iain J. Marshall, Thomas A. Trikalinos +1
Medical systematic reviews play a vital role in healthcare decision making and policy. However, their production is time-consuming, limiting the availability of high-quality and up…
Jointly Extracting Interventions, Outcomes, and Findings from RCT Reports with LLMs
Somin Wadhwa, Jay DeYoung, Benjamin Nye +2
Results from Randomized Controlled Trials (RCTs) establish the comparative effectiveness of interventions, and are in turn critical inputs for evidence-based care. However, results…