Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
arXiv:2408.07665 · doi:10.1109/SLT61566.2024.10832259
Abstract
Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit biases due to the nature of their training data. Recently, more Speech Large Language Models (SLLMs) have emerged, underscoring the urgent need to address these biases. This study introduces Spoken Stereoset, a dataset specifically designed to evaluate social biases in SLLMs. By examining how different models respond to speech from diverse demographic groups, we aim to identify these biases. Our experiments reveal significant insights into their performance and bias levels. The findings indicate that while most models show minimal bias, some still exhibit slightly stereotypical or anti-stereotypical tendencies.
References in corpus (10)
- NExT-GPT: Any-to-Any Multimodal LLM
- Quantifying Bias in Automatic Speech Recognition
- Bias in Automated Speaker Recognition
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
- Cognitive Bias in Decision-Making with LLMs
- Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
- Bias patterns in the application of LLMs for clinical decision support: A comprehensive study
- On the social bias of speech self-supervised models