2 citations · 7 across the 8 of their papers we have counts for
8 papers
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
Fei Deng, Qifei Wang, Wei Wei +2
Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using…
Binaural Angular Separation Network
Yang Yang, George Sung, Shao-Fu Shih +3
We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with sim…
StreamVC: Real-Time Low-Latency Voice Conversion
Yang Yang, Yury Kartynnik, Yunpeng Li +4
We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlik…
On-device Real-time Custom Hand Gesture Recognition
Esha Uboweja, David Tian, Qifei Wang +5
Most existing hand gesture recognition (HGR) systems are limited to a predefined set of gestures. However, users and developers often want to recognize new, unseen gestures. This i…
Blendshapes GHUM: Real-time Monocular Facial Blendshape Prediction
Ivan Grishchenko, Geng Yan, Eduard Gabriel Bazavan +5
We present Blendshapes GHUM, an on-device ML pipeline that predicts 52 facial blendshape coefficients at 30+ FPS on modern mobile phones, from a single monocular RGB image and enab…
Towards Authentic Face Restoration with Iterative Diffusion Models and Beyond
Yang Zhao, Tingbo Hou, Yu-Chuan Su +2
An authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Mos…