4 papers
Music Tagging with Classifier Group Chains
Takuya Hasumi, Tatsuya Komatsu, Yusuke Fujita
We propose music tagging with classifier chains that model the interplay of music tags. Most conventional methods estimate multiple tags independently by treating them as multiple…
Audio Fingerprinting with Holographic Reduced Representations
Yusuke Fujita, Tatsuya Komatsu
This paper proposes an audio fingerprinting model with holographic reduced representation (HRR). The proposed method reduces the number of stored fingerprints, whereas conventional…
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
Michael Hentschel, Yuta Nishikawa, Tatsuya Komatsu +1
This study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil t…
Audio Difference Learning for Audio Captioning
Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda +1
This study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a f…