papers

Publications (18)

eess.AS2023

Neural Model Reprogramming with Similarity Based Mapping for Low-Resource Spoken Command Recognition

Hao Yen, Pin-Jui Ku, Chao-Han Huck Yang +4

In this study, we propose a novel adversarial reprogramming (AR) approach for low-resource spoken command recognition (SCR), and build an AR-SCR system. The AR procedure aims to mo…

cs.CV2025

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

Xinsong Zhang, Yarong Zeng, Xinting Huang +4

In recent years, the field of vision-language model pre-training has experienced rapid advancements, driven primarily by the continuous enhancement of textual capabilities in large…

eess.AS2024

Bayesian adaptive learning to latent variables via Variational Bayes and Maximum a Posteriori

Hu Hu, Sabato Marco Siniscalchi, Chin-Hui Lee

In this work, we aim to establish a Bayesian adaptive learning framework by focusing on estimating latent variables in deep neural network (DNN) models. Latent variables indeed enc…

eess.AS2020

An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances

Hu Hu, Sabato Marco Siniscalchi, Yannan Wang +3

In this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (…

cs.CV2023

TeachCLIP: Multi-Grained Teaching for Efficient Text-to-Video Retrieval

Kaibin Tian, Ruixiang Zhao, Hu Hu +4

For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods are dominating. Compared to CLIP4Clip which is efficient a…

cs.SD2022

A Lottery Ticket Hypothesis Framework for Low-Complexity Device-Robust Neural Acoustic Scene Classification

Hao Yen, Chao-Han Huck Yang, Hu Hu +9

We propose a novel neural model compression strategy combining data augmentation, knowledge transfer, pruning, and quantization for device-robust acoustic scene classification (ASC…

cs.SD2020

A Two-Stage Approach to Device-Robust Acoustic Scene Classification

Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13

To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convol…

cs.CL2020

Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition

Hu Hu, Rui Zhao, Jinyu Li +2

Recently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of bei…

eess.AS2025

Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer

Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang +1

In this work, we propose a novel variational Bayesian adaptive learning approach for cross-domain knowledge transfer to address acoustic mismatches between training and testing con…

eess.AS2022

A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer

Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang +1

We propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address ac…

eess.AS2021

REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling

Hu Hu, Xuesong Yang, Zeynab Raeesy +6

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…

eess.AS2020

Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification

Hu Hu, Sabato Marco Siniscalchi, Yannan Wang +1

In this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relat…

cs.CL2019

Improving RNN Transducer Modeling for End-to-End Speech Recognition

Jinyu Li, Rui Zhao, Hu Hu +1

In the last few years, an emerging trend in automatic speech recognition research is the study of end-to-end (E2E) systems. Connectionist Temporal Classification (CTC), Attention E…

cs.MM2022

A study on joint modeling and data augmentation of multi-modalities for audio-visual scene classification

Qing Wang, Jun Du, Siyuan Zheng +8

In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-tr…

eess.AS2020

L-Vector: Neural Label Embedding for Domain Adaptation

Zhong Meng, Hu Hu, Jinyu Li +4

We propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target dom…

eess.AS2020

Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement

Jun Qi, Hu Hu, Yannan Wang +3

This paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for s…

eess.AS2020

Tensor-to-Vector Regression for Multi-channel Speech Enhancement based on Tensor-Train Network

Jun Qi, Hu Hu, Yannan Wang +3

We propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key…

eess.AS2020

Device-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation

Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13

In this technical report, we present a joint effort of four groups, namely GT, USTC, Tencent, and UKE, to tackle Task 1 - Acoustic Scene Classification (ASC) in the DCASE 2020 Chal…