activity
20162023
most citedPointCAT: Cross-Attention Transformer for point cloud

6 citations · 10 across the 15 of their papers we have counts for

collaborators

15 papers

cs.CL2023

Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization

Luyao Cheng, Siqi Zheng, Zhang Qinglin +3

Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization ap…

eess.AS2023

CASA-ASR: Context-Aware Speaker-Attributed ASR

Mohan Shi, Zhihao Du, Qian Chen +5

Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…

eess.AS20231 cited

Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction

Mohan Shi, Yuchun Shu, Lingyun Zuo +4

For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach…

cs.CV20236 cited

PointCAT: Cross-Attention Transformer for point cloud

Xincheng Yang, Mingze Jin, Weiji He +1

Transformer-based models have significantly advanced natural language processing and computer vision in recent years. However, due to the irregular and disordered structure of poin…

math.RA2023

Images of linear polynomials on upper triangular matrix algebras

Yingyu Luo, Qian Chen

The Fagundes-Mello conjecture asserts that every multilinear polynomial on upper triangular matrix algebras is a vector space, which is an improtant variation of the old and famous…

cs.CL20231 cited

MUG: A General Meeting Understanding and Generation Benchmark

Qinglin Zhang, Chong Deng, Jiaqing Liu +7

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings…