activity
20182026
most citedIntroducing v0.5 of the AI Safety Benchmark from MLCommons

5 citations · 5 across the 5 of their papers we have counts for

collaborators

6 papers

cs.HC2026

"Always Want to Use it for Everything": Understanding Young Adults' Perceptions of AI Dependence

Ashlee Milton, Leah Ajmani, Amy Heger +3

The growing integration of general-purpose AI chatbots into people's daily lives has raised concerns about the potential for unhealthy dependence, particularly among young adults.…

cs.AI2026

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

Jina Suh, Mihaela Vorvoreanu, Forough Poursabzi-Sangdeh +5

As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist…

cs.CR2026

Media Integrity and Authentication: Status, Directions, and Futures

Jessica Young, Sam Vaughan, Andrew Jenks +9

We provide background on emerging challenges and future directions with media integrity and authentication methods, focusing on distinguishing AI-generated media from authentic con…

cs.CL2024★ 5 cited

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…

cs.SE2022

Aligning Offline Metrics and Human Judgments of Value for Code Generation Models

Victor Dibia, Adam Fourney, Gagan Bansal +3

Large language models have demonstrated great potential to assist programmers in generating code. For such human-AI pair programming scenarios, we empirically demonstrate that whil…

cs.AI2018

Manipulating and Measuring Model Interpretability

Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman +2

With machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Altho…