Publications (15)
Large Language Models, and LLM-Based Agents, Should Be Used to Enhance the Digital Public Sphere
Seth Lazar, Luke Thorburn, Tian Jin +1
This paper argues that large language model-based recommenders can displace today's attention-allocation machinery. LLM-based recommenders would ingest open-web content, infer a us…
The 2021 RecSys Challenge Dataset: Fairness is not optional
Luca Belli, Alykhan Tejani, Frank Portman +10
After the success the RecSys 2020 Challenge, we are describing a novel and bigger dataset that was released in conjunction with the ACM RecSys Challenge 2021. This year's dataset i…
From Optimizing Engagement to Measuring Value
Smitha Milli, Luca Belli, Moritz Hardt
Most recommendation engines today are based on predicting user engagement, e.g. predicting whether a user will click on an item or not. However, there is potentially a large gap be…
Random Isn't Always Fair: Candidate Set Imbalance and Exposure Inequality in Recommender Systems
Amanda Bower, Kristian Lum, Tomo Lazovich +2
Traditionally, recommender systems operate by returning a user a set of items, ranked in order of estimated relevance to that user. In recent years, methods relying on stochastic o…
A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by English Marginal Abuse Models on Twitter
Kyra Yee, Alice Schoenauer Sebag, Olivia Redfield +3
Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such dispropor…
VERA-MH Concept Paper
Luca Belli, Kate H. Bentley, Will Alexander +5
We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initi…
Measuring Disparate Outcomes of Content Recommendation Algorithms with Distributional Inequality Metrics
Tomo Lazovich, Luca Belli, Aaron Gonzales +5
The harmful impacts of algorithmic decision systems have recently come into focus, with many examples of systems such as machine learning (ML) models amplifying existing societal b…
County-level Algorithmic Audit of Racial Bias in Twitter's Home Timeline
Luca Belli, Kyra Yee, Uthaipon Tantipongpipat +3
We report on the outcome of an audit of Twitter's Home Timeline ranking system. The goal of the audit was to determine if authors from some racial groups experience systematically…
Causal Inference Struggles with Agency on Online Platforms
Smitha Milli, Luca Belli, Moritz Hardt
Online platforms regularly conduct randomized experiments to understand how changes to the platform causally affect various outcomes of interest. However, experimentation on online…
VERA-MH: Validation of Ethical and Responsible AI in Mental Health
Luca Belli, Kate H. Bentley, Josh Gieringer +6
Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Validations of Ethical and Respon…
Algorithmic Amplification of Politics on Twitter
Ferenc Huszár, Sofia Ira Ktena, Conor O'Brien +3
Content on Twitter's home timeline is selected and ordered by personalization algorithms. By consistently ranking certain content higher, these algorithms may amplify some messages…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Kate H. Bentley, Luca Belli, Adam M. Chekroud +7
Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are sa…
Privacy-Aware Recommender Systems Challenge on Twitter's Home Timeline
Luca Belli, Sofia Ira Ktena, Alykhan Tejani +12
Recommender systems constitute the core engine of most social network platforms nowadays, aiming to maximize user satisfaction along with other key business objectives. Twitter is…
Assessing Demographic Bias in Named Entity Recognition
Shubhanshu Mishra, Sijun He, Luca Belli
Named Entity Recognition (NER) is often the first step towards automated Knowledge Base (KB) generation from raw text. In this work, we assess the bias in various Named Entity Reco…
Fighting Redundancy and Model Decay with Embeddings
Dan Shiebler, Luca Belli, Jay Baxter +2
Every day, hundreds of millions of new Tweets containing over 40 languages of ever-shifting vernacular flow through Twitter. Models that attempt to extract insight from this fireho…