6 papers
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11
While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
Shomik Jain, Jack Lanchantin, Maximilian Nickel +4
Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of pro…
Representative Ranking for Deliberation in the Public Sphere
Manon Revel, Smitha Milli, Tyler Lu +2
Online comment sections, such as those on news sites or social media, have the potential to foster informal public deliberation, However, this potential is often undermined by the…
Algorithmic Fairness and Color-blind Racism: Navigating the Intersection
Jamelle Watson-Daniels
Our focus lies at the intersection between two broader research perspectives: (1) the scientific study of algorithms and (2) the scholarship on race and racism. Many streams of res…
Predictive Churn with the Set of Good Models
Jamelle Watson-Daniels, Flavio du Pin Calmon, Alexander D'Amour +3
Issues can arise when research focused on fairness, transparency, or safety is conducted separately from research driven by practical deployment concerns and vice versa. This separ…