works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.SD2026

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

Xilin Jiang, Riki Shimizu, Sukru Samet Dindar +3

The paper presents Cocktail-Talker, a speech‑language model framework that lets a spoken assistant decide whether to respond, listen, or ignore in multi‑speaker, noisy social conve…

cs.SD2026

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

Sukru Samet Dindar, Riki Shimizu, Xilin Jiang +1

Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To ad…

cs.SD2026

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Xilin Jiang, Qiaolin Wang, Junkai Wu +30

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…

eess.AS2025

MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow

Riki Shimizu, Xilin Jiang, Nima Mesgarani

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent ad…

cs.SD2025

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Qiaolin Wang, Xilin Jiang, Linyang He +2

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-…

cs.CL2025

Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations

Linyang He, Qiaolin Wang, Xilin Jiang +1

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode s…