12 papers
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +36
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…
Disentangling Geometry, Performance, and Training in Language Models
Atharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian +1
Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estima…
Challenges to Grassroots Organization Engagement with AI Policy
Carter Buckner, Jennifer Mickel, Nandhini Swaminathan +6
Public policies are being developed around the world to address privacy, economic, intellectual property, energy, and other risks that AI technologies pose. Involvement from the ge…
Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends
Sabine Weber, Angelina Wang, Ankush Gupta +16
Natural language processing (NLP) technologies are rapidly reshaping how language is created, processed, and interpreted by humans. With current and potential applications in hirin…
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11
While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…
Centering Ecological Goals in Automated Identification of Individual Animals
Lukas Picek, Timm Haucke, Lukáš Adam +16
Recognizing individual animals over time is central to many ecological and conservation questions, including estimating abundance, survival, movement, and social structure. Recent…