papers

Publications (15)

cs.CY2025

Large Language Models, and LLM-Based Agents, Should Be Used to Enhance the Digital Public Sphere

Seth Lazar, Luke Thorburn, Tian Jin +1

This paper argues that large language model-based recommenders can displace today's attention-allocation machinery. LLM-based recommenders would ingest open-web content, infer a us…

cs.SI2021

The 2021 RecSys Challenge Dataset: Fairness is not optional

Luca Belli, Alykhan Tejani, Frank Portman +10

After the success the RecSys 2020 Challenge, we are describing a novel and bigger dataset that was released in conjunction with the ACM RecSys Challenge 2021. This year's dataset i…

cs.SI2021

From Optimizing Engagement to Measuring Value

Smitha Milli, Luca Belli, Moritz Hardt

Most recommendation engines today are based on predicting user engagement, e.g. predicting whether a user will click on an item or not. However, there is potentially a large gap be…

cs.IR2022

Random Isn't Always Fair: Candidate Set Imbalance and Exposure Inequality in Recommender Systems

Amanda Bower, Kristian Lum, Tomo Lazovich +2

Traditionally, recommender systems operate by returning a user a set of items, ranked in order of estimated relevance to that user. In recent years, methods relying on stochastic o…

cs.CL2022

A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by English Marginal Abuse Models on Twitter

Kyra Yee, Alice Schoenauer Sebag, Olivia Redfield +3

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such dispropor…

cs.CY2026

VERA-MH Concept Paper

Luca Belli, Kate H. Bentley, Will Alexander +5

We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initi…

cs.CY2022

Measuring Disparate Outcomes of Content Recommendation Algorithms with Distributional Inequality Metrics

Tomo Lazovich, Luca Belli, Aaron Gonzales +5

The harmful impacts of algorithmic decision systems have recently come into focus, with many examples of systems such as machine learning (ML) models amplifying existing societal b…

cs.SI2023

County-level Algorithmic Audit of Racial Bias in Twitter's Home Timeline

Luca Belli, Kyra Yee, Uthaipon Tantipongpipat +3

We report on the outcome of an audit of Twitter's Home Timeline ranking system. The goal of the audit was to determine if authors from some racial groups experience systematically…

cs.LG2022

Causal Inference Struggles with Agency on Online Platforms

Smitha Milli, Luca Belli, Moritz Hardt

Online platforms regularly conduct randomized experiments to understand how changes to the platform causally affect various outcomes of interest. However, experimentation on online…

cs.AI2026

VERA-MH: Validation of Ethical and Responsible AI in Mental Health

Luca Belli, Kate H. Bentley, Josh Gieringer +6

Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Validations of Ethical and Respon…

cs.CY2021

Algorithmic Amplification of Politics on Twitter

Ferenc Huszár, Sofia Ira Ktena, Conor O'Brien +3

Content on Twitter's home timeline is selected and ordered by personalization algorithms. By consistently ranking certain content higher, these algorithms may amplify some messages…

cs.AI2026

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Kate H. Bentley, Luca Belli, Adam M. Chekroud +7

Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are sa…

cs.SI2020

Privacy-Aware Recommender Systems Challenge on Twitter's Home Timeline

Luca Belli, Sofia Ira Ktena, Alykhan Tejani +12

Recommender systems constitute the core engine of most social network platforms nowadays, aiming to maximize user satisfaction along with other key business objectives. Twitter is…

cs.CL2020

Assessing Demographic Bias in Named Entity Recognition

Shubhanshu Mishra, Sijun He, Luca Belli

Named Entity Recognition (NER) is often the first step towards automated Knowledge Base (KB) generation from raw text. In this work, we assess the bias in various Named Entity Reco…

cs.SI2018

Fighting Redundancy and Model Decay with Embeddings

Dan Shiebler, Luca Belli, Jay Baxter +2

Every day, hundreds of millions of new Tweets containing over 40 languages of ever-shifting vernacular flow through Twitter. Models that attempt to extract insight from this fireho…