papers

Publications (54)

eess.AS2026

Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi +2

To advance immersive communication, the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge recently introduced Task 4 on Spatial Semantic Segmentatio…

cs.LG2025

FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local Parameters

Hiro Ishii, Kenta Niwa, Hiroshi Sawada +3

We propose Federated Preconditioned Mixing (FedPM), a novel Federated Learning (FL) method that leverages second-order optimization. Prior methods--such as LocalNewton, LTDA, and F…

eess.AS2026

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

Yanzhou Ren, Noboru Harada, Daiki Takeuchi +6

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensurin…

eess.AS2020

The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi +2

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captionin…

eess.AS2022

BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Pre-trained models are essential as feature extractors in modern machine learning systems in various domains. In this study, we hypothesize that representations effective for gener…

eess.AS2019

Data-driven design of perfect reconstruction filterbank for DNN-based sound source enhancement

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to est…

eess.AS2021

Description and Discussion on DCASE 2021 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions

Yohei Kawaguchi, Keisuke Imoto, Yuma Koizumi +6

We present the task description and discussion on the results of the DCASE 2021 Challenge Task 2. In 2020, we organized an unsupervised anomalous sound detection (ASD) task, identi…

eess.AS2021

ToyADMOS2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions

Noboru Harada, Daisuke Niizumi, Daiki Takeuchi +3

This paper proposes a new large-scale dataset called "ToyADMOS2" for anomaly detection in machine operating sounds (ADMOS). As did for our previous ToyADMOS dataset, we collected a…

cs.SD2022

Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques

Kota Dohi, Keisuke Imoto, Noboru Harada +7

We present the task description and discussion on the results of the DCASE 2022 Challenge Task 2: ``Unsupervised anomalous sound detection (ASD) for machine condition monitoring ap…

eess.SP2023

Deep sound-field denoiser: optically-measured sound-field denoising using deep neural network

Kenji Ishikawa, Daiki Takeuchi, Noboru Harada +1

This paper proposes a deep sound-field denoiser, a deep neural network (DNN) based denoising of optically measured sound-field images. Sound-field imaging using optical methods has…

eess.AS2022

Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval

Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi +2

The amount of audio data available on public websites is growing rapidly, and an efficient mechanism for accessing the desired data is necessary. We propose a content-based audio r…

eess.AS2024

Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN

Shiqi Zhang, Zheng Qiu, Daiki Takeuchi +2

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhanc…

cs.SD2025

Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Masahiro Yasuda, Binh Thien Nguyen, Noboru Harada +10

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound ev…

cs.SD2019

Deep Griffin-Lim Iteration

Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi +2

This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). T…

eess.AS2022

Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this stud…

eess.AS2019

Trainable Adaptive Window Switching for Speech Enhancement

Yuma Koizumi, Noboru Harada, Yoichi Haneda

This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform…

cs.SD2025

Acousto-optic reconstruction of exterior sound field based on concentric circle sampling with circular harmonic expansion

Phuc Duc Nguyen, Kenji Ishikawa, Noboru Harada +1

Acousto-optic sensing provides an alternative approach to traditional microphone arrays by shedding light on the interaction of light with an acoustic field. Sound field reconstruc…

eess.AS2024

M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +4

Contrastive language-audio pre-training (CLAP) enables zero-shot (ZS) inference of audio and exhibits promising performance in several classification tasks. However, conventional a…

eess.AS2022

Masked Spectrogram Modeling using Masked Autoencoders for Learning General-purpose Audio Representation

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Recent general-purpose audio representations show state-of-the-art performance on various audio tasks. These representations are pre-trained by self-supervised learning methods tha…

eess.AS2026

Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Tomoya Nishida, Noboru Harada, Daiki Takeuchi +6

This paper presents an overview of DCASE 2026 Challenge Task 2, titled "Noise-aware unsupervised anomalous sound detection (UASD) for machine condition monitoring." The task aims t…

eess.AS2025

Towards Pre-training an Effective Respiratory Audio Foundation Model

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to d…

eess.AS2020

Invertible DNN-based nonlinear time-frequency transform for speech enhancement

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose an end-to-end speech enhancement method with trainable time-frequency~(T-F) transform based on invertible deep neural network~(DNN). The resent development of speech enh…

eess.AS2019

ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection

Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu +2

This paper introduces a new dataset called "ToyADMOS" designed for anomaly detection in machine operating sounds (ADMOS). To the best our knowledge, no large-scale datasets are ava…

eess.AS2023

Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods…

eess.AS2023

Masked Modeling Duo for Speech: Specializing General-Purpose Audio Representation to Speech using Denoising Distillation

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Self-supervised learning general-purpose audio representations have demonstrated high performance in a variety of tasks. Although they can be optimized for application by fine-tuni…

eess.AS2019

First Order Ambisonics Domain Spatial Augmentation for DNN-based Direction of Arrival Estimation

Luca Mazzon, Yuma Koizumi, Masahiro Yasuda +1

In this paper, we propose a novel data augmentation method for training neural networks for Direction of Arrival (DOA) estimation. This method focuses on expanding the representati…

eess.SP2024

SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes

Risako Tanigawa, Kenji Ishikawa, Noboru Harada +1

Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acousto-optic sensing enables understanding of the interaction between sound and ob…

stat.ML2019

AdaFlow: Domain-Adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-Domain Translation

Masataka Yamaguchi, Yuma Koizumi, Noboru Harada

We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recentl…

eess.AS2026

Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation…

eess.AS2024

Description and Discussion on DCASE 2024 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Tomoya Nishida, Noboru Harada, Daisuke Niizumi +9

We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge Task 2: First-shot unsupervised anomalous sound detection (…

eess.AS2026

Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Binh Thien Nguyen, Masahiro Yasuda, Noboru Harada +8

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5).…

cs.SD2025

Description and Discussion on DCASE 2025 Challenge Task 2: First-shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Tomoya Nishida, Noboru Harada, Daisuke Niizumi +9

This paper introduces the task description for the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge Task 2, titled "First-shot unsupervised anomalo…

eess.AS2020

Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto +8

In this paper, we present the task description and discuss the results of the DCASE 2020 Challenge Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitori…

eess.AS2019

Batch Uniformization for Minimizing Maximum Anomaly Score of DNN-based Anomaly Detection in Sounds

Yuma Koizumi, Shoichiro Saito, Masataka Yamaguchi +2

Use of an autoencoder (AE) as a normal model is a state-of-the-art technique for unsupervised-anomaly detection in sounds (ADS). The AE is trained to minimize the sample mean of th…

eess.AS2024

Masked Modeling Duo: Towards a Universal Audio Pre-training Framework

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved ma…

eess.AS2024

Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval

Shunsuke Tsubaki, Daisuke Niizumi, Daiki Takeuchi +3

The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech…

eess.AS2025

Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi +3

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the D…

eess.AS2020

Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning

Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2

The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…

eess.AS2023

Audio Difference Captioning Utilizing Similarity-Discrepancy Disentanglement

Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi +2

We proposed Audio Difference Captioning (ADC) as a new extension task of audio captioning for describing the semantic differences between input pairs of similar but slightly differ…

eess.AS2020

Phase reconstruction based on recurrent phase unwrapping with deep neural networks

Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi +2

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio s…

eess.AS2025

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda +3

Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhan…

eess.AS2025

Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domain…

eess.AS2022

Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion

Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito +1

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microph…

eess.AS2024

Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

To reduce the need for skilled clinicians in heart sound interpretation, recent studies on automating cardiac auscultation have explored deep learning approaches. However, despite…

eess.AS2022

ConceptBeam: Concept Driven Target Speech Extraction

Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai +6

We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speake…

eess.AS2020

Real-time speech enhancement using equilibriated RNN

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose a speech enhancement method using a causal deep neural network~(DNN) for real-time applications. DNN has been widely used for estimating a time-frequency~(T-F) mask whic…

eess.AS2025

M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Contrastive language-audio pre-training (CLAP), which learns audio-language representations by aligning audio and text in a common feature space, has become popular for solving aud…

eess.AS2021

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio represen…

cs.MM2024

Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis

Masahiro Yasuda, Noboru Harada, Yasunori Ohishi +3

Observations with distributed sensors are essential in analyzing a series of human and machine activities (referred to as 'events' in this paper) in complex and extensive real-worl…

eess.AS2024

6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human

Masahiro Yasuda, Shoichiro Saito, Akira Nakayama +1

We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with micr…

cs.SD2023

Description and Discussion on DCASE 2023 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Kota Dohi, Keisuke Imoto, Noboru Harada +7

We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge Task 2: ``First-shot unsupervised anomalous sound detection…

cs.SD2025

Learning to assess subjective impressions from speech

Yuto Kondo, Hirokazu Kameoka, Kou Tanaka +2

We tackle a new task of training neural network models that can assess subjective impressions conveyed through speech and assign scores accordingly, inspired by the work on automat…

eess.AS2023

First-shot anomaly sound detection for machine condition monitoring: A domain generalization baseline

Noboru Harada, Daisuke Niizumi, Yasunori Ohishi +2

This paper provides a baseline system for First-shot-compliant unsupervised anomaly detection (ASD) for machine condition monitoring. First-shot ASD does not allow systems to do ma…

stat.ML2018

Unsupervised Detection of Anomalous Sound based on Deep Learning and the Neyman-Pearson Lemma

Yuma Koizumi, Shoichiro Saito, Hisashi Uematsum Yuta Kawachi +1

This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of unsupervised-ADS…