14 citations · 19 across the 5 of their papers we have counts for
11 papers
When Incentives Backfire, Data Stops Being Human
Sebastin Santy, Prasanta Bhattacharya, Manoel Horta Ribeiro +2
Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality…
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
Rock Yuren Pang, Sebastin Santy, René Just +1
Digital technologies have positively transformed society, but they have also led to undesirable consequences not anticipated at the time of design or development. We posit that ins…
Multilingual Diversity Improves Vision-Language Representations
Thao Nguyen, Matthew Wallingford, Sebastin Santy +5
Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on s…
Semantic and Expressive Variation in Image Captions Across Languages
Andre Ye, Sebastin Santy, Jena D. Hwang +2
Computer vision often treats human perception as homogeneous: an implicit assumption that visual stimuli are perceived similarly by everyone. This assumption is reflected in the wa…
NLPositionality: Characterizing Design Biases of Datasets and Models
Sebastin Santy, Jenny T. Liang, Ronan Le Bras +2
Design biases in NLP systems, such as performance differences for different populations, often stem from their creator's positionality, i.e., views and lived experiences shaped by…
Learnings from Technological Interventions in a Low Resource Language: Enhancing Information Access in Gondi
Devansh Mehta, Harshita Diddee, Ananya Saxena +8
The primary obstacle to developing technologies for low-resource languages is the lack of representative, usable data. In this paper, we report the deployment of technology-driven…