collaborators

6 papers

cs.CL2025

Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence

Sophia Hager, David Mueller, Kevin Duh +1

As large language models (LLMs) are increasingly used for factual question-answering, it becomes more important for LLMs to have the capability to communicate the likelihood that t…

cs.CL2025

Whisper-UT: A Unified Translation Framework for Speech and Text

Cihan Xiao, Matthew Wiesner, Debashish Chakraborty +7

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challen…

eess.AS2025

GenVC: Self-Supervised Zero-Shot Voice Conversion

Zexin Cai, Henry Li Xinyuan, Ashi Garg +5

Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this d…

eess.AS2025

Scalable Controllable Accented TTS

Henry Li Xinyuan, Zexin Cai, Ashi Garg +5

We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for…

eess.AS2025

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

Ashi Garg, Zexin Cai, Lin Zhang +6

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what…

cs.CL2024

SpeechQE: Estimating the Quality of Direct Speech Translation

HyoJung Han, Kevin Duh, Marine Carpuat

Recent advances in automatic quality estimation for machine translation have exclusively focused on written language, leaving the speech modality underexplored. In this work, we fo…