3 papers
cs.CL2026
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
Keren Fuentes, Aaron Mueller
Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have…
cs.LG2025
Privacy-Preserving Dataset Combination
Keren Fuentes, Mimee Xu, Irene Chen
Access to diverse, high-quality datasets is crucial for machine learning model performance, yet data sharing remains limited by privacy concerns and competitive interests, particul…
cs.CL2022
Language model acceptability judgements are not always robust to context
Koustuv Sinha, Jon Gauthier, Aaron Mueller +4
Targeted syntactic evaluations of language models ask whether models show stable preferences for syntactically acceptable content over minimal-pair unacceptable inputs. Most target…