4 papers · 1 filter
Disentangling Geometry, Performance, and Training in Language Models
Atharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian +1
Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estima…
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
Christina Chance, Rebecca Pattichis, Arjun Subramonian +4
Reclaimed slur usage is a common and meaningful practice online for many marginalized communities. It serves as a source of solidarity, identity, and shared experience. However, co…
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
Arjun Subramonian, Vagrant Gautam, Preethi Seshadri +3
Numerous methods have been proposed to measure LLM misgendering, including probability-based evaluations (e.g., automatically with templatic sentences) and generation-based evaluat…
Understanding "Democratization" in NLP and ML Research
Arjun Subramonian, Vagrant Gautam, Dietrich Klakow +1
Recent improvements in natural language processing (NLP) and machine learning (ML) and increased mainstream adoption have led to researchers frequently discussing the "democratizat…