3 papers
cs.SD2026
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Xilin Jiang, Qiaolin Wang, Junkai Wu +30
Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…
eess.AS2025
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
Riki Shimizu, Xilin Jiang, Nima Mesgarani
Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent ad…
q-bio.NC2025
Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models
Riki Shimizu, Richard J. Antonello, Chandan Singh +1
Speech foundation models (SFMs) are increasingly hailed as powerful computational models of human speech perception. However, since their representations are inherently black-box,…