Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
arXiv:2009.11436
Abstract
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, dataaugmentation, multi-task learning, and post-processing, for audiocaptioning. The system received the highest evaluation scores, butwhich of the individual elements most fully contributed to its perfor-mance has not yet been clarified. Here, to asses their contributions,we first conducted an element-wise ablation study on our systemto estimate to what extent each element is effective. We then con-ducted a detailed module-wise ablation study to further clarify thekey processing modules for improving accuracy. The results showthat data augmentation and post-processing significantly improvethe score in our system. In particular, mix-up data augmentationand beam search in post-processing improve SPIDEr by 0.8 and 1.6points, respectively.
Accepted to DCASE2020 Workshop
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- Unsupervised Data Augmentation for Consistency Training
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- Acoustic Scene Classification
- EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
- Augmenting Data with Mixup for Sentence Classification: An Empirical Study
- Pre-training via Paraphrasing
- Audio Caption: Listen and Tell
- The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation
- Neural Machine Translation For Paraphrase Generation
Cited by in corpus (5)
- Automated Audio Captioning: An Overview of Recent Progress and New Challenges
- Audio Captioning Transformer
- An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
- CL4AC: A Contrastive Loss for Audio Captioning
- WaveTransformer: A Novel Architecture for Audio Captioning Based on Learning Temporal and Time-Frequency Information