activity
20232025
most citedChanging Answer Order Can Decrease MMLU Accuracy

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

Bhaktipriya Radharapu, Manon Revel, Megan Ung +2

The increasing use of LLMs as substitutes for humans in ``aligning'' LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambiv…

cs.CL2024

Chained Tuning Leads to Biased Forgetting

Megan Ung, Alicia Sun, Samuel J. Bell +3

Large language models (LLMs) are often fine-tuned for use on downstream tasks, though this can degrade capabilities learned during previous training. This phenomenon, often referre…

cs.CL2024

Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Vipul Gupta, Candace Ross, David Pantoja +3

One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the qual…

cs.CL20241 cited

Changing Answer Order Can Decrease MMLU Accuracy

Vipul Gupta, David Pantoja, Candace Ross +2

As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. M…

cs.CL2023

ROBBIE: Robust Bias Evaluation of Large Generative Language Models

David Esiobu, Xiaoqing Tan, Saghar Hosseini +7

As generative large language models (LLMs) grow more performant and prevalent, we must develop comprehensive enough tools to measure and improve their fairness. Different prompt-ba…