6 citations · 10 across the 15 of their papers we have counts for
15 papers
Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization
Luyao Cheng, Siqi Zheng, Zhang Qinglin +3
Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization ap…
CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen +5
Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo +4
For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach…
PointCAT: Cross-Attention Transformer for point cloud
Xincheng Yang, Mingze Jin, Weiji He +1
Transformer-based models have significantly advanced natural language processing and computer vision in recent years. However, due to the irregular and disordered structure of poin…
Images of linear polynomials on upper triangular matrix algebras
Yingyu Luo, Qian Chen
The Fagundes-Mello conjecture asserts that every multilinear polynomial on upper triangular matrix algebras is a vector space, which is an improtant variation of the old and famous…
MUG: A General Meeting Understanding and Generation Benchmark
Qinglin Zhang, Chong Deng, Jiaqing Liu +7
Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings…