papers

Publications (5)

eess.AS2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…

eess.AS2022

TorchAudio: Building Blocks for Audio and Speech Processing

Yao-Yuan Yang, Moto Hira, Zhaoheng Ni +20

This document describes version 0.10 of TorchAudio: building blocks for machine learning applications in the audio and speech processing domain. The objective of TorchAudio is to a…

cs.IT2014

Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks

Junjie Zhu, Siddhartan Govindasamy, Jeff Hwang

A technique is presented to characterize the Signal-to-Interference-plus-Noise Ratio (SINR) of a representative link with a multiantenna linear Minimum-Mean-Square-Error receiver i…

eess.AS2023

TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch

Jeff Hwang, Moto Hira, Caroline Chen +21

TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing…

cs.SD2025

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Xueyao Zhang, Xiaohui Zhang, Kainan Peng +10

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotat…