5 citations · 11 across the 3 of their papers we have counts for
9 papers
Grapheme-to-Phoneme Transformer Model for Transfer Learning Dialects
Eric Engelhart, Mahsa Elyasi, Gaurav Bharaj
Grapheme-to-Phoneme (G2P) models convert words to their phonetic pronunciations. Classic G2P methods include rule-based systems and pronunciation dictionaries, while modern G2P sys…
Generative Landmarks
David Ferman, Gaurav Bharaj
We propose a general purpose approach to detect landmarks with improved temporal consistency, and personalization. Most sparse landmark detection methods rely on laborious, manuall…
Flavored Tacotron: Conditional Learning for Prosodic-linguistic Features
Mahsa Elyasi, Gaurav Bharaj
Neural sequence-to-sequence text-to-speech synthesis (TTS), such as Tacotron-2, transforms text into high-quality speech. However, generating speech with natural prosody still rema…
Generalized Spoofing Detection Inspired from Audio Generation Artifacts
Yang Gao, Tyler Vuong, Mahsa Elyasi +2
State-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains. Such artifacts could be well capt…
Practical Face Reconstruction via Differentiable Ray Tracing
Abdallah Dib, Gaurav Bharaj, Junghyun Ahn +4
We present a differentiable ray-tracing based novel face reconstruction approach where scene attributes - 3D geometry, reflectance (diffuse, specular and roughness), pose, camera p…
StyleRig: Rigging StyleGAN for 3D Control over Portrait Images
Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj +5
StyleGAN generates photorealistic portrait images of faces with eyes, teeth, hair and context (neck, shoulders, background), but lacks a rig-like control over semantic face paramet…