1 citations · 1 across the 2 of their papers we have counts for
10 papers
VIRTUE: Visual-Interactive Text-Image Universal Embedder
Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu +2
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embe…
Diffusion-based Signal Refiner for Speech Enhancement and Separation
Masato Hirano, Ryosuke Sawata, Naoki Murata +2
Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper propos…
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
Shuichiro Nishigori, Koichi Saito, Naoki Murata +3
Speech enhancement (SE) utilizing diffusion models is a promising technology that improves speech quality in noisy speech data. Furthermore, the Schrödinger bridge (SB) has recent…
LOCKEY: A Novel Approach to Model Authentication and Deepfake Tracking
Mayank Kumar Singh, Naoya Takahashi, Wei-Hsiang Liao +1
This paper presents a novel approach to deter unauthorized deepfakes and enable user tracking in generative models, even when the user has full access to the model parameters, by i…
SilentCipher: Deep Audio Watermarking
Mayank Kumar Singh, Naoya Takahashi, Weihsiang Liao +1
In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advanceme…
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
Sungho Lee, Marco MartÃnez-RamÃrez, Wei-Hsiang Liao +4
We present GRAFX, an open-source library designed for handling audio processing graphs in PyTorch. Along with various library functionalities, we describe technical details on the…