Publications (18)
Neural Model Reprogramming with Similarity Based Mapping for Low-Resource Spoken Command Recognition
Hao Yen, Pin-Jui Ku, Chao-Han Huck Yang +4
In this study, we propose a novel adversarial reprogramming (AR) approach for low-resource spoken command recognition (SCR), and build an AR-SCR system. The AR procedure aims to mo…
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
Xinsong Zhang, Yarong Zeng, Xinting Huang +4
In recent years, the field of vision-language model pre-training has experienced rapid advancements, driven primarily by the continuous enhancement of textual capabilities in large…
Bayesian adaptive learning to latent variables via Variational Bayes and Maximum a Posteriori
Hu Hu, Sabato Marco Siniscalchi, Chin-Hui Lee
In this work, we aim to establish a Bayesian adaptive learning framework by focusing on estimating latent variables in deep neural network (DNN) models. Latent variables indeed enc…
An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances
Hu Hu, Sabato Marco Siniscalchi, Yannan Wang +3
In this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (…
TeachCLIP: Multi-Grained Teaching for Efficient Text-to-Video Retrieval
Kaibin Tian, Ruixiang Zhao, Hu Hu +4
For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods are dominating. Compared to CLIP4Clip which is efficient a…
A Lottery Ticket Hypothesis Framework for Low-Complexity Device-Robust Neural Acoustic Scene Classification
Hao Yen, Chao-Han Huck Yang, Hu Hu +9
We propose a novel neural model compression strategy combining data augmentation, knowledge transfer, pruning, and quantization for device-robust acoustic scene classification (ASC…
A Two-Stage Approach to Device-Robust Acoustic Scene Classification
Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13
To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convol…
Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition
Hu Hu, Rui Zhao, Jinyu Li +2
Recently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of bei…
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang +1
In this work, we propose a novel variational Bayesian adaptive learning approach for cross-domain knowledge transfer to address acoustic mismatches between training and testing con…
A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer
Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang +1
We propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address ac…
REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling
Hu Hu, Xuesong Yang, Zeynab Raeesy +6
Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…
Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification
Hu Hu, Sabato Marco Siniscalchi, Yannan Wang +1
In this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relat…
Improving RNN Transducer Modeling for End-to-End Speech Recognition
Jinyu Li, Rui Zhao, Hu Hu +1
In the last few years, an emerging trend in automatic speech recognition research is the study of end-to-end (E2E) systems. Connectionist Temporal Classification (CTC), Attention E…
A study on joint modeling and data augmentation of multi-modalities for audio-visual scene classification
Qing Wang, Jun Du, Siyuan Zheng +8
In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-tr…
L-Vector: Neural Label Embedding for Domain Adaptation
Zhong Meng, Hu Hu, Jinyu Li +4
We propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target dom…
Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement
Jun Qi, Hu Hu, Yannan Wang +3
This paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for s…
Tensor-to-Vector Regression for Multi-channel Speech Enhancement based on Tensor-Train Network
Jun Qi, Hu Hu, Yannan Wang +3
We propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key…
Device-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation
Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13
In this technical report, we present a joint effort of four groups, namely GT, USTC, Tencent, and UKE, to tackle Task 1 - Acoustic Scene Classification (ASC) in the DCASE 2020 Chal…