activity
20242026
most citedWavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convolution and Harmonic Prior for Reliable Complex Spectrogram Estimation

1 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

17 papers · 1 filter

eess.AS2026

Pseudo-label distillation for discriminative anomalous sound detection

Takuya Fujimura, Tomoki Toda

Discriminative anomalous sound detection (ASD) methods train a feature extractor through a classification task using machine-information labels. They then detect anomalies in the r…

eess.AS2026

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

Ding Ma, Jinyi Mi, Fengji Li +5

Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, lim…

eess.AS2025

Handling Domain Shifts for Anomalous Sound Detection: A Review of DCASE-Related Work

Kevin Wilkinghoff, Takuya Fujimura, Keisuke Imoto +3

When detecting anomalous sounds in complex environments, one of the main difficulties is that trained models must be sensitive to subtle differences in monitored target signals, wh…

eess.AS2025

Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake

Liping Chen, Kong Aik Lee, Zhen-Hua Ling +4

In the era of big data, remarkable advancements have been achieved in personalized speech generation techniques that utilize speaker attributes, including voice and speaking style,…

eess.AS2025

Layer-wise Analysis for Quality of Multilingual Synthesized Speech

Erica Cooper, Takuma Okamoto, Yamato Ohtani +2

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders t…

eess.AS2025

Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition

Cheng-Hung Hu, Yusuke Yasuda, Akifumi Yoshimoto +1

Speech Quality Assessment (SQA) and Continuous Speech Emotion Recognition (CSER) are two key tasks in speech technology, both relying on listener ratings. However, these ratings ar…