168 citations · 646 across the 40 of their papers we have counts for
4 papers · 2 filters
Nudging: Inference-time Alignment of LLMs via Guided Decoding
Yu Fei, Yasaman Razeghi, Sameer Singh
Large language models (LLMs) require alignment to effectively and safely follow user instructions. This process necessitates training an aligned version for every base model, resul…
Perceptions of Linguistic Uncertainty by Language Models and Humans
Catarina G Belem, Markelle Kelly, Mark Steyvers +2
_Uncertainty expressions_ such as "probably" or "highly unlikely" are pervasive in human language. While prior work has established that there is population-level agreement in term…
Are Models Biased on Text without Gender-related Language?
Catarina G Belém, Preethi Seshadri, Yasaman Razeghi +1
Gender bias research has been pivotal in revealing undesirable behaviors in large language models, exposing serious gender stereotypes associated with occupations, and emotions. A…
MisgenderMender: A Community-Informed Approach to Interventions for Misgendering
Tamanna Hossain, Sunipa Dev, Sameer Singh
Content Warning: This paper contains examples of misgendering and erasure that could be offensive and potentially triggering. Misgendering, the act of incorrectly addressing someon…