activity
20222026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

Xilin Jiang, Riki Shimizu, Sukru Samet Dindar +3

Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, a…

cs.SD2026

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Xilin Jiang, Qiaolin Wang, Junkai Wu +30

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…

cs.SD2025

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Qiaolin Wang, Xilin Jiang, Linyang He +2

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-…

cs.SD2025

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

Xilin Jiang, Junkai Wu, Vishal Choudhari +1

Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-…

cs.SD2023

Meta-AF Echo Cancellation for Improved Keyword Spotting

Jonah Casebeer, Junkai Wu, Paris Smaragdis

Adaptive filters (AFs) are vital for enhancing the performance of downstream tasks, such as speech recognition, sound event detection, and keyword spotting. However, traditional AF…

cs.SD2023

Unsupervised Improvement of Audio-Text Cross-Modal Representations

Zhepei Wang, Cem Subakan, Krishna Subramani +4

Recent advances in using language models to obtain cross-modal audio-text representations have overcome the limitations of conventional training approaches that use predefined labe…