activity
20242026
collaborators

7 papers

cs.SD2026

Do speech foundation models perceive speaker similarity as humans do?

Minoru Kishi, Hayato Yagi, Shinnosuke Takamichi +1

This study presents a comparative analysis between the speaker embeddings of speech foundation models and human subjective perception of speaker similarity. Human listeners have th…

cs.CL2026

Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation

Ryota Kawamatsu, Anum Afzal, Yuki Saito +5

We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is…

eess.AS2026

Human-CLAP: Human-perception-based contrastive language-audio pretraining

Taisei Takano, Yuki Okamoto, Yusuke Kanamori +3

Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, h…

cs.CL2026

Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches

Anum Afzal, Yuki Saito, Hiroya Takamura +5

Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and li…

cs.SD2025

Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions

Kentaro Seki, Yuki Okamoto, Kouei Yamaoka +3

Contrastive language--audio pretraining (CLAP) has achieved remarkable success as an audio--text embedding framework, but existing approaches are limited to monaural or single-sour…

cs.SD2025

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Yusuke Kanamori, Yuki Okamoto, Taisei Takano +3

In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and…