1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 1 cited
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
Luyao Cheng, Hui Wang, Siqi Zheng +5
Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the inter…
cs.CL2024
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
Qian Chen, Wen Wang, Qinglin Zhang +7
The Transformer architecture has significantly advanced deep learning, particularly in natural language processing, by effectively managing long-range dependencies. However, as the…