5 papers
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Harshita Chopra, Kshitish Ghate, Aylin Caliskan +3
Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are unclear, impatient, or reluctant…
Characterizing Resource Sharing Practices on Underground Internet Forum Synthetic Non-Consensual Intimate Image Content Creation Communities
Bernardo B. P. Medeiros, Malvika Jadhav, Allison Lu +3
Many malicious actors responsible for disseminating synthetic non-consensual intimate imagery (SNCII) operate within internet forums to exchange resources, strategies, and generate…
Examining Risks Through a Characterization of the AI Companion Application Ecosystem: A Stratified Sample from the Apple App Store and Google Play Store
Natalie Grace Brigham, Lucy Qin, Tadayoshi Kohno
While computer systems that allow users to interact through conversational natural language (i.e., chatbots) have existed for many years, various types of applications offering AI…
A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
Rachel Hong, Jevan Hutson, William Agnew +3
We investigate the contents of web-scraped data for training AI systems, at sizes where human dataset curators and compilers no longer manually annotate every sample. Building off…
Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
Rachel Hong, William Agnew, Tadayoshi Kohno +1
As training datasets become increasingly drawn from unstructured, uncontrolled environments such as the web, researchers and industry practitioners have increasingly relied upon da…