activity
20242026
most citedSpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays

1 citations · 1 across the 10 of their papers we have counts for

collaborators

12 papers

quant-ph2026

Optimal pure state cloning and transposition are complementary channels

Vanessa Brzić, Dmitry Grinko, Michał Studziński +1

State cloning and state transposition are fundamental transformations which, despite being desirable, cannot be perfectly realised due to two conceptually distinct constraints of q…

eess.AS20261 cited

SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays

Yiwen Shao, Yong Xu, Sanjeev Khudanpur +1

Spatial information is a critical clue for multi-channel multi-speaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems ext…

cs.CL2026

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

Kalvin Chang, Yiwen Shao, Jiahong Li +1

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Ma…

cs.CL2025

AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning

Yiwen Shao, Wei Liu, Jiahong Li +4

Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a…

eess.AS2025

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao +2

Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward…

eess.AS2025

TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation

Wei Liu, Jiahong Li, Yiwen Shao +1

Speech-LLM models have demonstrated great performance in multi-modal and multi-task speech understanding. A typical speech-LLM paradigm is integrating speech modality with a large…