Automated Audio Captioning: An Overview of Recent Progress and New Challenges
arXiv:2205.05949 · doi:10.1186/s13636-022-00259-2
Abstract
Automated audio captioning is a cross-modal translation task that aims to generate natural language descriptions for given audio clips. This task has received increasing attention with the release of freely available datasets in recent years. The problem has been addressed predominantly with deep learning techniques. Numerous approaches have been proposed, such as investigating different neural network architectures, exploiting auxiliary information such as keywords or sentence information to guide caption generation, and employing different training strategies, which have greatly facilitated the development of this field. In this paper, we present a comprehensive review of the published contributions in automated audio captioning, from a variety of existing approaches to evaluation metrics and datasets. We also discuss open challenges and envisage possible future research directions.
Accepted by EURASIP Journal on Audio Speech and Music Processing
References in corpus (17)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Sequence Transduction with Recurrent Neural Networks
- Acoustic Scene Classification
- Sound Event Detection: A Tutorial
- Audio Retrieval with Natural Language Queries: A Benchmark Study
- Automated Audio Captioning: An Overview of Recent Progress and New Challenges
- Audio Captioning Transformer
- Masked Non-Autoregressive Image Captioning
- An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
- Multi-task Regularization Based on Infrequent Classes for Audio Captioning
- Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
- Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
- Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning
- Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning
- Evaluating Off-the-Shelf Machine Listening and Natural Language Models for Automated Audio Captioning
- Continual Learning for Automated Audio Captioning Using The Learning Without Forgetting Approach
Cited by in corpus (5)
- WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
- Automated Audio Captioning: An Overview of Recent Progress and New Challenges
- Graph Attention for Automated Audio Captioning
- Audio-Language Datasets of Scenes and Events: A Survey
- Towards Generating Diverse Audio Captions via Adversarial Training