From the 1 of 13 linked papers with an AI index.
13 papers
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
Xilin Jiang, Riki Shimizu, Sukru Samet Dindar +3
The paper presents Cocktail-Talker, a speech‑language model framework that lets a spoken assistant decide whether to respond, listen, or ignore in multi‑speaker, noisy social conve…
Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning
Sukru Samet Dindar, Riki Shimizu, Xilin Jiang +1
Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To ad…
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Xilin Jiang, Qiaolin Wang, Junkai Wu +30
Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
Riki Shimizu, Xilin Jiang, Nima Mesgarani
Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent ad…
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
Qiaolin Wang, Xilin Jiang, Linyang He +2
While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-…
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
Linyang He, Qiaolin Wang, Xilin Jiang +1
Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode s…