Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
arXiv:2407.06957 · doi:10.1109/SLT61566.2024.10832317
Abstract
Speech Integrated Large Language Models (SILLMs) combine large language models with speech perception to perform diverse tasks, such as emotion recognition to speaker verification, demonstrating universal audio understanding capability. However, these models may amplify biases present in training data, potentially leading to biased access to information for marginalized groups. This work introduces a curated spoken bias evaluation toolkit and corresponding dataset. We evaluate gender bias in SILLMs across four semantic-related tasks: speech-to-text translation (STT), spoken coreference resolution (SCR), spoken sentence continuation (SSC), and spoken question answering (SQA). Our analysis reveals that bias levels are language-dependent and vary with different evaluation methods. Our findings emphasize the necessity of employing multiple approaches to comprehensively assess biases in SILLMs, providing insights for developing fairer SILLM systems.
References in corpus (11)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Gender bias and stereotypes in Large Language Models
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- Quantifying Bias in Automatic Speech Recognition
- Seamless: Multilingual Expressive and Streaming Speech Translation
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
- On the social bias of speech self-supervised models
- CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models