collaborators

5 papers

cs.SD2025

Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Masahiro Yasuda, Binh Thien Nguyen, Noboru Harada +10

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound ev…

eess.AS2025

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda +3

Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhan…

eess.AS2025

Towards Pre-training an Effective Respiratory Audio Foundation Model

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to d…

eess.AS2025

Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domain…

eess.AS2025

Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi +3

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the D…