activity
20192024
most citedBOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation

208 citations · 226 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2024

Fewer Truncations Improve Language Modeling

Hantian Ding, Zijian Wang, Giovanni Paolini +4

In large language model training, input documents are typically concatenated together and then split into sequences of equal length to avoid padding tokens. Despite its efficiency,…

cs.CL20225 cited

Is the Elephant Flying? Resolving Ambiguities in Text-to-Image Generative Models

Ninareh Mehrabi, Palash Goyal, Apurv Verma +7

Natural language often contains ambiguities that can lead to misinterpretation and miscommunication. While humans can handle ambiguities effectively by asking clarifying questions…

cs.CL2022

An Analysis of the Effects of Decoding Algorithms on Fairness in Open-Ended Language Generation

Jwala Dhamala, Varun Kumar, Rahul Gupta +2

Several prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in deter…

cs.CL202210 cited

On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations

Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang +4

Multiple metrics have been introduced to measure fairness in various natural language processing tasks. These metrics can be roughly categorized into two categories: 1) \emph{extri…

cs.CL20222 cited

Mitigating Gender Bias in Distilled Language Models via Counterfactual Role Reversal

Umang Gupta, Jwala Dhamala, Varun Kumar +7

Language models excel at generating coherent text, and model compression techniques such as knowledge distillation have enabled their use in resource-constrained settings. However,…

cs.CL2021

Industry Scale Semi-Supervised Learning for Natural Language Understanding

Luoxin Chen, Francisco Garcia, Varun Kumar +2

This paper presents a production Semi-Supervised Learning (SSL) pipeline based on the student-teacher framework, which leverages millions of unlabeled examples to improve Natural L…