123 citations · 233 across the 18 of their papers we have counts for
18 papers
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
Henry Li Xinyuan, Zexin Cai, Ashi Garg +5
We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-t…
Clean Label Attacks against SLU Systems
Henry Li Xinyuan, Sonal Joshi, Thomas Thebaud +3
Poisoning backdoor attacks involve an adversary manipulating the training data to induce certain behaviors in the victim model by inserting a trigger in the signal at inference tim…
Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
Zexin Cai, Henry Li Xinyuan, Ashi Garg +5
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has…
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
Ruizhe Huang, Mahsa Yarmohammadi, Sanjeev Khudanpur +1
Existing research suggests that automatic speech recognition (ASR) models can benefit from additional contexts (e.g., contact lists, user specified vocabulary). Rare words and name…
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shi-Xiong Zhang, Yong Xu +4
In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background nois…