2 papers
cs.SD2026
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
Ayuto Tsutsumi, Kohei Tanaka, Sayaka Shiota
In this paper, we propose a submission to the x-to-audio alignment (XACLE) challenge. The goal is to predict semantic alignment of a given general audio and text pair. The proposed…
eess.AS2025
Voice Privacy Preservation with Multiple Random Orthogonal Secret Keys: Attack Resistance Analysis
Kohei Tanaka, Hitoshi Kiya, Sayaka Shiota
Recently, opportunities to transmit speech data to deep learning models executed in the cloud have increased. This has led to growing concerns about speech privacy, including both…