activity
20172023
most citedAutoencoder Regularized Network For Driving Style Representation Learning

4 citations · 4 across the 3 of their papers we have counts for

collaborators

10 papers

eess.AS2024

VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Yuke Lin, Ming Cheng, Fulin Zhang +3

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This…

cs.CL2024

Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification

Yaqian Hao, Chenguang Hu, Yingying Gao +2

The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-di…

eess.AS2024

On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations

Yaqian Hao, Chenguang Hu, Yingying Gao +2

For speech classification tasks, deep learning models often achieve high accuracy but exhibit shortcomings in calibration, manifesting as classifiers exhibiting overconfidence. The…

eess.AS2024

CEC: A Noisy Label Detection Method for Speaker Recognition

Yao Shen, Yingying Gao, Yaqian Hao +4

Noisy labels are inevitable, even in well-annotated datasets. The detection of noisy labels is of significant importance to enhance the robustness of speaker recognition models. In…

eess.AS2024

Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network

Yanan Chen, Zihao Cui, Yingying Gao +3

The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due…

cs.LG2023

Cascaded Multi-task Adaptive Learning Based on Neural Architecture Search

Yingying Gao, Shilei Zhang, Zihao Cui +2

Cascading multiple pre-trained models is an effective way to compose an end-to-end system. However, fine-tuning the full cascaded model is parameter and memory inefficient and our…