5 papers
Harnessing Textual Refusal Directions for Multimodal Safety
Moreno D'IncÃ, Nicu Sebe, Massimiliano Mancini
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feas…
Safe Vision-Language Models via Unsafe Weights Manipulation
Moreno D'IncÃ, Elia Peruzzo, Xingqian Xu +3
Vision-language models (VLMs) often inherit the biases and unsafe associations present within their large-scale training dataset. While recent approaches mitigate unsafe behaviors,…
Socially Pertinent Robots in Gerontological Healthcare
Xavier Alameda-Pineda, Angus Addlesee, Daniel Hernández GarcÃa +41
Despite the many recent achievements in developing and deploying social robotics, there are still many underexplored environments and applications for which systematic evaluation o…
Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models
Aditya Gulati, Moreno D'IncÃ, Nicu Sebe +2
Physical attractiveness matters. It has been shown to influence human perception and decision-making, often leading to biased judgments that favor those deemed attractive in what i…
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
Quentin Guimard, Moreno D'IncÃ, Massimiliano Mancini +1
A person downloading a pre-trained model from the web should be aware of its biases. Existing approaches for bias identification rely on datasets containing labels for the task of…