8 citations · 9 across the 15 of their papers we have counts for
7 papers · 1 filter
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
Miriam Wanner, Benjamin Van Durme, Mark Dredze
The decompose-then-verify strategy for verification of Large Language Model (LLM) generations decomposes claims that are then independently verified. Decontextualization augments t…
Are Clinical T5 Models Better for Clinical Text?
Yahan Li, Keith Harrigian, Ayah Zirikly +1
Large language models with a transformer-based encoder/decoder architecture, such as T5, have become standard platforms for supervised tasks. To bring these technologies to the cli…
Give me Some Hard Questions: Synthetic Data Generation for Clinical QA
Fan Bai, Keith Harrigian, Joel Stremmel +3
Clinical Question Answering (QA) systems enable doctors to quickly access patient information from electronic health records (EHRs). However, training these systems requires signif…
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Sharon Levy, William D. Adler, Tahilin Sanchez Karver +2
Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated mo…
Transferring Fairness using Multi-Task Learning with Limited Demographic Information
Carlos Aguirre, Mark Dredze
Training supervised machine learning systems with a fairness loss can improve prediction fairness across different demographic groups. However, doing so requires demographic annota…
A Closer Look at Claim Decomposition
Miriam Wanner, Seth Ebner, Zhengping Jiang +2
As generated text becomes more commonplace, it is increasingly important to evaluate how well-supported such text is by external knowledge sources. Many approaches for evaluating t…