activity
20242026
most citedBenchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

8 citations · 9 across the 15 of their papers we have counts for

collaborators
Showing 2024Show all

7 papers · 1 filter

cs.CL2024

DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation

Miriam Wanner, Benjamin Van Durme, Mark Dredze

The decompose-then-verify strategy for verification of Large Language Model (LLM) generations decomposes claims that are then independently verified. Decontextualization augments t…

cs.CL2024

Are Clinical T5 Models Better for Clinical Text?

Yahan Li, Keith Harrigian, Ayah Zirikly +1

Large language models with a transformer-based encoder/decoder architecture, such as T5, have become standard platforms for supervised tasks. To bring these technologies to the cli…

cs.CL2024

Give me Some Hard Questions: Synthetic Data Generation for Clinical QA

Fan Bai, Keith Harrigian, Joel Stremmel +3

Clinical Question Answering (QA) systems enable doctors to quickly access patient information from electronic health records (EHRs). However, training these systems requires signif…

cs.CL2024

Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Sharon Levy, William D. Adler, Tahilin Sanchez Karver +2

Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated mo…

cs.LG2024

Transferring Fairness using Multi-Task Learning with Limited Demographic Information

Carlos Aguirre, Mark Dredze

Training supervised machine learning systems with a fairness loss can improve prediction fairness across different demographic groups. However, doing so requires demographic annota…

cs.CL2024

A Closer Look at Claim Decomposition

Miriam Wanner, Seth Ebner, Zhengping Jiang +2

As generated text becomes more commonplace, it is increasingly important to evaluate how well-supported such text is by external knowledge sources. Many approaches for evaluating t…