Publications (35)
Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources
Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi +2
To advance immersive communication, the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge recently introduced Task 4 on Spatial Semantic Segmentatio…
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
Yanzhou Ren, Noboru Harada, Daiki Takeuchi +6
Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensurin…
The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation
Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi +2
This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captionin…
BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Pre-trained models are essential as feature extractors in modern machine learning systems in various domains. In this study, we hypothesize that representations effective for gener…
Data-driven design of perfect reconstruction filterbank for DNN-based sound source enhancement
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2
We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to est…
ToyADMOS2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions
Noboru Harada, Daisuke Niizumi, Daiki Takeuchi +3
This paper proposes a new large-scale dataset called "ToyADMOS2" for anomaly detection in machine operating sounds (ADMOS). As did for our previous ToyADMOS dataset, we collected a…
Deep sound-field denoiser: optically-measured sound-field denoising using deep neural network
Kenji Ishikawa, Daiki Takeuchi, Noboru Harada +1
This paper proposes a deep sound-field denoiser, a deep neural network (DNN) based denoising of optically measured sound-field images. Sound-field imaging using optical methods has…
Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval
Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi +2
The amount of audio data available on public websites is growing rapidly, and an efficient mechanism for accessing the desired data is necessary. We propose a content-based audio r…
Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
Shiqi Zhang, Zheng Qiu, Daiki Takeuchi +2
With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhanc…
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
Masahiro Yasuda, Binh Thien Nguyen, Noboru Harada +10
Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound ev…
Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this stud…
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +4
Contrastive language-audio pre-training (CLAP) enables zero-shot (ZS) inference of audio and exhibits promising performance in several classification tasks. However, conventional a…
Masked Spectrogram Modeling using Masked Autoencoders for Learning General-purpose Audio Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Recent general-purpose audio representations show state-of-the-art performance on various audio tasks. These representations are pre-trained by self-supervised learning methods tha…
Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi +2
The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to th…
Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Tomoya Nishida, Noboru Harada, Daiki Takeuchi +6
This paper presents an overview of DCASE 2026 Challenge Task 2, titled "Noise-aware unsupervised anomalous sound detection (UASD) for machine condition monitoring." The task aims t…
Towards Pre-training an Effective Respiratory Audio Foundation Model
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to d…
Invertible DNN-based nonlinear time-frequency transform for speech enhancement
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2
We propose an end-to-end speech enhancement method with trainable time-frequency~(T-F) transform based on invertible deep neural network~(DNN). The resent development of speech enh…
Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods…
Masked Modeling Duo for Speech: Specializing General-Purpose Audio Representation to Speech using Denoising Distillation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Self-supervised learning general-purpose audio representations have demonstrated high performance in a variety of tasks. Although they can be optimized for application by fine-tuni…
Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention
Yuma Koizumi, Kohei Yatabe, Marc Delcroix +2
This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly fro…
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation…
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
Binh Thien Nguyen, Masahiro Yasuda, Noboru Harada +8
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5).…
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved ma…
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
Shunsuke Tsubaki, Daisuke Niizumi, Daiki Takeuchi +3
The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech…
Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi +3
Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the D…
Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…
Audio Difference Captioning Utilizing Similarity-Discrepancy Disentanglement
Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi +2
We proposed Audio Difference Captioning (ADC) as a new extension task of audio captioning for describing the semantic differences between input pairs of similar but slightly differ…
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda +3
Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhan…
Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domain…
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
To reduce the need for skilled clinicians in heart sound interpretation, recent studies on automating cardiac auscultation have explored deep learning approaches. However, despite…
ConceptBeam: Concept Driven Target Speech Extraction
Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai +6
We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speake…
Real-time speech enhancement using equilibriated RNN
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2
We propose a speech enhancement method using a causal deep neural network~(DNN) for real-time applications. DNN has been widely used for estimating a time-frequency~(T-F) mask whic…
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Contrastive language-audio pre-training (CLAP), which learns audio-language representations by aligning audio and text in a common feature space, has become popular for solving aud…
BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio represen…
First-shot anomaly sound detection for machine condition monitoring: A domain generalization baseline
Noboru Harada, Daisuke Niizumi, Yasunori Ohishi +2
This paper provides a baseline system for First-shot-compliant unsupervised anomaly detection (ASD) for machine condition monitoring. First-shot ASD does not allow systems to do ma…