2 papers
cs.SD2026
Do speech foundation models perceive speaker similarity as humans do?
Minoru Kishi, Hayato Yagi, Shinnosuke Takamichi +1
This study presents a comparative analysis between the speaker embeddings of speech foundation models and human subjective perception of speaker similarity. Human listeners have th…
cs.SD2025
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
Minoru Kishi, Ryosuke Sakai, Shinnosuke Takamichi +2
We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the s…