37 papers
Evaluating Large Language Models for Antisemitic Incident Classification
Karina Halevy, Julia Mendelsohn, Chan Young Park +2
Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We…
The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor
Jordan Taylor, William Agnew, Maarten Sap +2
Visual generative AI models are trained using a one-size-fits-all measure of aesthetic appeal. However, what is deemed "aesthetic" is inextricably linked to personal taste and cult…
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim +2
Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all…
Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
Joel Mire, Maria Antoniak, Steven R. Wilson +4
Reading stories evokes rich interpretive, affective, and evaluative responses, such as inferences about narrative intent or judgments about characters. Yet, computational models of…
Olmo 3
Team Olmo, :, Allyson Ettinger +66
We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long-context reasoning, function…
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
Sanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola +2
AI agents are increasingly being deployed to automate tasks, often based on underspecified user instructions. Making unwarranted assumptions to compensate for the missing informati…