activity
20162025
most citedEnsemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection

19 citations · 106 across the 23 of their papers we have counts for

collaborators

32 papers

cs.SD2025★ 1 cited

Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification

Kazuki Shimada, Archontis Politis, Iran R. Roman +10

This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the c…

cs.CR2024

LOCKEY: A Novel Approach to Model Authentication and Deepfake Tracking

Mayank Kumar Singh, Naoya Takahashi, Wei-Hsiang Liao +1

This paper presents a novel approach to deter unauthorized deepfakes and enable user tracking in generative models, even when the user has full access to the model parameters, by i…

cs.SD2024★ 18 cited

SilentCipher: Deep Audio Watermarking

Mayank Kumar Singh, Naoya Takahashi, Weihsiang Liao +1

In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advanceme…

cs.SD2023★ 7 cited

STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam +9

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perce…

cs.SD2023

Iteratively Improving Speech Recognition and Voice Conversion

Mayank Kumar Singh, Naoya Takahashi, Onoe Naoyuki

Many existing works on voice conversion (VC) tasks use automatic speech recognition (ASR) models for ensuring linguistic consistency between source and converted samples. However,…

eess.AS2023★ 3 cited

The Whole Is Greater than the Sum of Its Parts: Improving Music Source Separation by Bridging Network

Ryosuke Sawata, Naoya Takahashi, Stefan Uhlich +2

This paper presents the crossing scheme (X-scheme) for improving the performance of deep neural network (DNN)-based music source separation (MSS) with almost no increasing calculat…