activity
20182024
most citedGraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network

19 citations · 33 across the 25 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD20231 cited

The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR

Yuhao Liang, Mohan Shi, Fan Yu +11

With the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle…

cs.SD2023

Profile-Error-Tolerant Target-Speaker Voice Activity Detection

Dongmei Wang, Xiong Xiao, Naoyuki Kanda +3

Target-Speaker Voice Activity Detection (TS-VAD) utilizes a set of speaker profiles alongside an input audio signal to perform speaker diarization. While its superiority over conve…

cs.SD20222 cited

The Microsoft System for VoxCeleb Speaker Recognition Challenge 2022

Gang Liu, Tianyan Zhou, Yong Zhao +4

In this report, we describe our submitted system for track 2 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). We fuse a variety of good-performing models ranging fro…

cs.SD2022

Maximizing Audio Event Detection Model Performance on Small Datasets Through Knowledge Transfer, Data Augmentation, And Pretraining: An Ablation Study

Daniel Tompkins, Kshitiz Kumar, Jian Wu

An Xception model reaches state-of-the-art (SOTA) accuracy on the ESC-50 dataset for audio event detection through knowledge transfer from ImageNet weights, pretraining on AudioSet…

cs.SD20202 cited

IEEE SLT 2021 Alpha-mini Speech Challenge: Open Datasets, Tracks, Rules and Baselines

Yihui Fu, Zhuoyuan Yao, Weipeng He +9

The IEEE Spoken Language Technology Workshop (SLT) 2021 Alpha-mini Speech Challenge (ASC) is intended to improve research on keyword spotting (KWS) and sound source location (SSL)…