4 papers
Multilingual Diversity Improves Vision-Language Representations
Thao Nguyen, Matthew Wallingford, Sebastin Santy +5
Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on s…
When Incentives Backfire, Data Stops Being Human
Sebastin Santy, Prasanta Bhattacharya, Manoel Horta Ribeiro +2
Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality…
Semantic and Expressive Variation in Image Captions Across Languages
Andre Ye, Sebastin Santy, Jena D. Hwang +2
Computer vision often treats human perception as homogeneous: an implicit assumption that visual stimuli are perceived similarly by everyone. This assumption is reflected in the wa…
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
Rock Yuren Pang, Sebastin Santy, René Just +1
Digital technologies have positively transformed society, but they have also led to undesirable consequences not anticipated at the time of design or development. We posit that ins…