From the 2 of 47 papers with an AI index.
18 citations
- Dartmouth HospitalGB16 papers
- California Institute of TechnologyUS7 papers
- Durham UniversityGB7 papers
- University of ChicagoUS7 papers
- Center for Astrophysics Harvard & SmithsonianUS6 papers
- Centre National de la Recherche ScientifiqueFR6 papers
- Goddard Space Flight CenterUS6 papers
- Instituto de Astrofísica de CanariasES6 papers
- Jet Propulsion LaboratoryUS6 papers
- University of Illinois Urbana-ChampaignUS6 papers
- Columbia UniversityUS5 papers
- Duke UniversityUS5 papers
3 papers · 1 filter
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation
Yahan Zheng, John Guerrerio, Soroush Vosoughi +1
Preprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP. While these approaches reduce measurable stereotypes…
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
Weicheng Ma, John Guerrerio, Soroush Vosoughi
Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual…
Spectral Signatures of Large Language Models
Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu +4
The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as mod…