Privacy-oriented manipulation of speaker representations
arXiv:2310.06652 · doi:10.1109/ACCESS.2024.3409067
Abstract
Speaker embeddings are ubiquitous, with applications ranging from speaker recognition and diarization to speech synthesis and voice anonymisation. The amount of information held by these embeddings lends them versatility, but also raises privacy concerns. Speaker embeddings have been shown to contain information on age, sex, health and more, which speakers may want to keep private, especially when this information is not required for the target task. In this work, we propose a method for removing and manipulating private attributes from speaker embeddings that leverages a Vector-Quantized Variational Autoencoder architecture, combined with an adversarial classifier and a novel mutual information loss. We validate our model on two attributes, sex and age, and perform experiments with ignorant and fully-informed attackers, and with in-domain and out-of-domain data.
Article published in IEEE Access
References in corpus (8)
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- SpeechBrain: A General-Purpose Speech Toolkit
- Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?
- Privacy-preserving Voice Analysis via Disentangled Representations
- A Step Towards Preserving Speakers' Identity While Detecting Depression Via Speaker Disentanglement
- The Privacy ZEBRA: Zero Evidence Biometric Recognition Assessment
- Joint gender and age estimation based on speech signals using x-vectors and transfer learning
- Beyond Neural-on-Neural Approaches to Speaker Gender Protection