26 citations · 189 across the 40 of their papers we have counts for
27 papers · 1 filter
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
Shuichiro Nishigori, Koichi Saito, Naoki Murata +3
Speech enhancement (SE) utilizing diffusion models is a promising technology that improves speech quality in noisy speech data. Furthermore, the Schrödinger bridge (SB) has recentl…
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
Sungho Lee, Marco Martínez-Ramírez, Wei-Hsiang Liao +4
We present GRAFX, an open-source library designed for handling audio processing graphs in PyTorch. Along with various library functionalities, we describe technical details on the…
SilentCipher: Deep Audio Watermarking
Mayank Kumar Singh, Naoya Takahashi, Weihsiang Liao +1
In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advanceme…
Zero- and Few-shot Sound Event Localization and Detection
Kazuki Shimada, Kengo Uchida, Yuichiro Koyama +4
Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems…
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
Takashi Shibuya, Yuhta Takida, Yuki Mitsufuji
Generative adversarial network (GAN)-based vocoders have been intensively studied because they can synthesize high-fidelity audio waveforms faster than real-time. However, it has b…
Automatic Piano Transcription with Hierarchical Frequency-Time Transformer
Keisuke Toyama, Taketo Akama, Yukara Ikemiya +3
Taking long-term spectral and temporal dependencies into account is essential for automatic piano transcription. This is especially helpful when determining the precise onset and o…