Publications (12)
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
Nanna Inie, Jonathan Stray, Leon Derczynski
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition o…
Building Human Values into Recommender Systems: An Interdisciplinary Synthesis
Jonathan Stray, Alon Halevy, Parisa Assar +18
Recommender systems are the algorithms which select, filter, and personalize content across many of the worlds largest platforms and apps. As such, their positive and negative effe…
Social Media Harm Abatement: Mechanisms for Transparent Public Health Assessment
Nathaniel Lubin, Yuning Liu, Amanda Yarnell +7
Social media platforms have been accused of causing a range of harms, resulting in dozens of lawsuits across jurisdictions. These lawsuits are situated within the context of a long…
The Prosocial Ranking Challenge: Reducing Polarization on Social Media without Sacrificing Engagement
Jonathan Stray, Ian Baker, George Beknazar-Yuzbashev +40
We report the first direct comparisons of multiple alternative social media algorithms on multiple platforms on outcomes of societal interest. We used a browser extension to modify…
What are you optimizing for? Aligning Recommender Systems with Human Values
Jonathan Stray, Ivan Vendrov, Jeremy Nixon +2
We describe cases where real recommender systems were modified in the service of various human values such as diversity, fairness, well-being, time well spent, and factual accuracy…
What We Know About Using Non-Engagement Signals in Content Ranking
Tom Cunningham, Sana Pandey, Leif Sigerson +7
Many online platforms predominantly rank items by predicted user engagement. We believe that there is much unrealized potential in including non-engagement signals, which can impro…