15 citations · 20 across the 7 of their papers we have counts for
3 papers · 1 filter
CHEF: Cross-modal Hierarchical Embeddings for Food Domain Retrieval
Hai X. Pham, Ricardo Guerrero, Jiatong Li +1
Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construc…
End-to-end Learning for 3D Facial Animation from Raw Waveforms of Speech
Hai X. Pham, Yuting Wang, Vladimir Pavlovic
We present a deep learning framework for real-time speech-driven 3D facial animation from just raw waveforms. Our deep neural network directly maps an input sequence of speech audi…
Robust Performance-driven 3D Face Tracking in Long Range Depth Scenes
Hai X. Pham, Chongyu Chen, Luc N. Dao +3
We introduce a novel robust hybrid 3D face tracking framework from RGBD video streams, which is capable of tracking head pose and facial actions without pre-calibration or interven…