collaborators

12 papers

eess.AS2026

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

Dehui Gao, Zhixian Zhao, Zhennan Lin +14

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentri…

eess.AS2026

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition

Naijun Zheng, Yuke Lin, Sanli Tian +4

Multi-talker speech recognition is often addressed by combining automatic speech recognition (ASR) and speaker diarization in a pipeline system. Recently, LLM-based approaches have…

eess.AS2026

EvoTSE: Evolving Enrollment for Target Speaker Extraction

Zikai Liu, Ziqian Wang, Xingchen Li +4

Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity…

eess.AS2026

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement

Xingchen Li, Hanke Xie, Ziqian Wang +4

Generative Universal Speech Enhancement (USE) methods aim to leverage generative models to improve speech quality under various types of distortions. However, existing generative s…

cs.SD2026

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

Zhixian Zhao, Shuiyuan Wang, Guojian Li +12

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowin…

cs.SD2025

LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation

Jun Chen, Shichao Hu, Jiuxin Lin +8

In-car multi-zone speech separation, which captures voices from different speech zones, plays a crucial role in human-vehicle interaction. Although previous SpatialNet has achieved…