collaborators

5 papers

cs.LG2024

Hypernetworks for Personalizing ASR to Atypical Speech

Max Müller-Eberstein, Dianna Yee, Karren Yang +2

Parameter-efficient fine-tuning (PEFT) for personalizing automatic speech recognition (ASR) has recently shown promise for adapting general population models to atypical speech. Ho…

cs.CV2023

FastSR-NeRF: Improving NeRF Efficiency on Consumer Devices with A Simple Super-Resolution Pipeline

Chien-Yu Lin, Qichen Fu, Thomas Merth +2

Super-resolution (SR) techniques have recently been proposed to upscale the outputs of neural radiance fields (NeRF) and generate high-quality images with enhanced inference speeds…

cs.CV2023

Probabilistic Speech-Driven 3D Facial Motion Synthesis: New Benchmarks, Methods, and Applications

Karren D. Yang, Anurag Ranjan, Jen-Hao Rick Chang +2

We consider the task of animating 3D facial geometry from speech signal. Existing works are primarily deterministic, focusing on learning a one-to-one mapping from speech signal to…

cs.SD2023

Novel-View Acoustic Synthesis from 3D Reconstructed Rooms

Byeongjoo Ahn, Karren Yang, Brian Hamilton +5

We investigate the benefit of combining blind audio recordings with 3D scene information for novel-view acoustic synthesis. Given audio recordings from 2-4 microphones and the 3D g…

eess.AS2023

Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models

Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula +5

While Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be finetuned on data…