9 papers
"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work
Anna Kawakami, Chloe Qianhui Zhao, Renee Shelby +3
Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluat…
"Death by a thousand taxonomies?": AI Risk Classification In Practice
Glen Berman, Ned Cooper, Angel Hsing-Chi Hwang +3
The harms in which AI is implicated range in nature and scope from unsafe user interactions through to the societal-wide consequences of AI adoption. Classification of the diverse…
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifica…
A Unified Framework to Quantify Cultural Intelligence of AI
Sunipa Dev, Vinodkumar Prabhakaran, Rutledge Chin Feman +16
As generative AI technologies are increasingly being launched across the globe, assessing their competence to operate in different cultural contexts is exigently becoming a priorit…
Cultural Perspectives and Expectations for Generative AI: A Global Survey Approach
Erin van Liemt, Renee Shelby, Andrew Smart +5
There is a lack of empirical evidence about global attitudes around whether and how GenAI should represent cultures. This paper assesses understandings and beliefs about culture as…
How Tech Workers Contend with Hazards of Humanlikeness in Generative AI
Mark DÃaz, Renee Shelby, Eric Corbett +1
Generative AI's humanlike qualities are driving its rapid adoption in professional domains. However, this anthropomorphic appeal raises concerns from HCI and responsible AI scholar…