The Ethical Implications of Generative Audio Models: A Systematic Literature Review
arXiv:2307.05527 · doi:10.1145/3600211.3604686
Abstract
Generative audio models typically focus their applications in music and speech generation, with recent models having human-like quality in their audio output. This paper conducts a systematic literature review of 884 papers in the area of generative audio models in order to both quantify the degree to which researchers in the field are considering potential negative impacts and identify the types of ethical implications researchers in this area need to consider. Though 65% of generative audio research papers note positive potential impacts of their work, less than 10% discuss any negative impacts. This jarringly small percentage of papers considering negative impact is particularly worrying because the issues brought to light by the few papers doing so are raising serious ethical implications and concerns relevant to the broader field such as the potential for fraud, deep-fakes, and copyright infringement. By quantifying this lack of ethical consideration in generative audio research and identifying key areas of potential harm, this paper lays the groundwork for future work in the field at a critical point in time in order to guide more conscientious research as this field progresses.
In proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES '23). 10 pages, 1 figure
References in corpus (70)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Showing Academic Performance Predictions during Term Planning: Effects on Students' Decisions, Behaviors, and Preferences
- Deep Learning for Audio Signal Processing
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- MusicLM: Generating Music From Text
- Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
- Quantifying Memorization Across Neural Language Models
- MelNet: A Generative Model for Audio in the Frequency Domain
- Transflower: probabilistic autoregressive dance generation with multimodal attention
- Universal Speech Enhancement with Score-based Diffusion
- AudioGen: Textually Guided Audio Generation
- CHiVE: Varying Prosody in Speech Synthesis with a Linguistically Driven Dynamic Hierarchical Conditional Variational Network
- Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
- BigVGAN: A Universal Neural Vocoder with Large-Scale Training
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation
- It's Time to Do Something: Mitigating the Negative Impacts of Computing Through a Change to the Peer Review Process
- The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism
- Unpacking the Expressed Consequences of AI Research in Broader Impact Statements
- Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction
- AI Song Contest: Human-AI Co-Creation in Songwriting
- Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data
- ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
- On Provable Copyright Protection for Generative Models
- Learning Interpretable Representation for Controllable Polyphonic Music Generation
- DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
- A Study on Speech Enhancement Based on Diffusion Probabilistic Model
- SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech
- Heterogeneous Target Speech Separation
- StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis
- Restoring degraded speech via a modified diffusion model
- Feature reinforcement with word embedding and parsing information in neural TTS
- Crowdsourcing Impacts: Exploring the Utility of Crowds for Anticipating Societal Impacts of Algorithmic Decision Making
- Explicitly Conditioned Melody Generation: A Case Study with Interdependent RNNs
- HooliGAN: Robust, High Quality Neural Vocoding
- NONOTO: A Model-agnostic Web Interface for Interactive Music Composition by Inpainting
- Generative Melody Composition with Human-in-the-Loop Bayesian Optimization
- Assisted Sound Sample Generation with Musical Conditioning in Adversarial Auto-Encoders
- The Piano Inpainting Application
- Effective parameter estimation methods for an ExcitNet model in generative text-to-speech systems
- Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech
- Parallel Synthesis for Autoregressive Speech Generation
- A Review of Intelligent Music Generation Systems
- Is Disentanglement enough? On Latent Representations for Controllable Music Generation
- MP3net: coherent, minute-long music generation from raw audio with a simple convolutional GAN
- Generative Modelling for Controllable Audio Synthesis of Expressive Piano Performance
- Energy Consumption of Deep Generative Audio Models
- Complex Recurrent Variational Autoencoder with Application to Speech Enhancement
- Incorporating Multi-Target in Multi-Stage Speech Enhancement Model for Better Generalization
- Challenges in creative generative models for music: a divergence maximization perspective
- Ethics and Creativity in Computer Vision
- Armor: A Benchmark for Meta-evaluation of Artificial Music
- Vertical-Horizontal Structured Attention for Generating Music with Chords
- Re-creation of Creations: A New Paradigm for Lyric-to-Melody Generation
- V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice Anonymization
- Expressive Communication: A Common Framework for Evaluating Developments in Generative Models and Steering Interfaces
- Building Synthetic Speaker Profiles in Text-to-Speech Systems
- A Laptop Ensemble Performance System using Recurrent Neural Networks
- HpRNet : Incorporating Residual Noise Modeling for Violin in a Variational Parametric Synthesizer
- Symbolic Music Loop Generation with Neural Discrete Representations
- Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models
- Music Generation with Temporal Structure Augmentation
- Time out of Mind: Generating Rate of Speech conditioned on emotion and speaker
- Conditional variational autoencoder to improve neural audio synthesis for polyphonic music sound
- NeuralDPS: Neural Deterministic Plus Stochastic Model with Multiband Excitation for Noise-Controllable Waveform Generation
- SinTra: Learning an inspiration model from a single multi-track music segment
- Comparision Of Adversarial And Non-Adversarial LSTM Music Generative Models