Computational bioacoustics with deep learning: a review and roadmap
arXiv:2112.06725 · doi:10.7717/peerj.13152
Abstract
Animal vocalisations and natural soundscapes are fascinating objects of study, and contain valuable evidence about animal behaviours, populations and ecosystems. They are studied in bioacoustics and ecoacoustics, with signal processing and analysis an important component. Computational bioacoustics has accelerated in recent decades due to the growth of affordable digital sound recording devices, and to huge progress in informatics such as big data, signal processing and machine learning. Methods are inherited from the wider field of deep learning, including speech and image processing. However, the tasks, demands and data characteristics are often different from those addressed in speech or music analysis. There remain unsolved problems, and tasks for which evidence is surely present in many acoustic signals, but not yet realised. In this paper I perform a review of the state of the art in deep learning for computational bioacoustics, aiming to clarify key concepts and identify and analyse knowledge gaps. Based on this, I offer a subjective but principled roadmap for computational bioacoustics with deep learning: topics that the community should aim to address, in order to make the most of future developments in AI and informatics, and to use audio data in answering zoological and ecological questions.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- WaveNet: A Generative Model for Raw Audio
- Attention-Based Models for Speech Recognition
- Scaling Laws for Neural Language Models
- Sound Event Detection: A Tutorial
- You Only Hear Once: A YOLO-like Algorithm for Audio Segmentation and Sound Event Detection
- Deep embedded clustering of coral reef bioacoustics
- Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes
- Long-distance Detection of Bioacoustic Events with Per-channel Energy Normalization
- Tiny Transformers for Environmental Sound Classification at the Edge
- Proposal-based Few-shot Sound Event Detection for Speech and Environmental Sounds with Perceivers
Cited by in corpus (15)
- Global birdsong embeddings enable superior transfer learning for bioacoustic classification
- Unsupervised classification to improve the quality of a bird song recording dataset
- Adaptive Representations of Sound for Automatic Insect Recognition
- pykanto: a python library to accelerate research on wild bird song
- A Bird Song Detector for improving bird identification through Deep Learning: a case study from Doñana
- Spectrogram features for audio and speech analysis
- Fitting Auditory Filterbanks with Multiresolution Neural Networks
- acoupi: An Open-Source Python Framework for Deploying Bioacoustic AI Models on Edge Devices
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network
- Exploring Meta Information for Audio-based Zero-shot Bird Classification
- Instabilities in Convnets for Raw Audio
- Detection and classification of vocal productions in large scale audio recordings
- Multi-layer attentive probing improves transfer of audio representations for bioacoustics
- Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice