activity
20222024
most citedEfficient Heterogeneous Video Segmentation at the Edge

2 citations · 7 across the 8 of their papers we have counts for

collaborators

8 papers

cs.LG20242 cited

PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models

Fei Deng, Qifei Wang, Wei Wei +2

Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using…

eess.AS2024

Binaural Angular Separation Network

Yang Yang, George Sung, Shao-Fu Shih +3

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with sim…

eess.AS20241 cited

StreamVC: Real-Time Low-Latency Voice Conversion

Yang Yang, Yury Kartynnik, Yunpeng Li +4

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlik…

cs.CV20231 cited

On-device Real-time Custom Hand Gesture Recognition

Esha Uboweja, David Tian, Qifei Wang +5

Most existing hand gesture recognition (HGR) systems are limited to a predefined set of gestures. However, users and developers often want to recognize new, unseen gestures. This i…

cs.CV20231 cited

Blendshapes GHUM: Real-time Monocular Facial Blendshape Prediction

Ivan Grishchenko, Geng Yan, Eduard Gabriel Bazavan +5

We present Blendshapes GHUM, an on-device ML pipeline that predicts 52 facial blendshape coefficients at 30+ FPS on modern mobile phones, from a single monocular RGB image and enab…

cs.CV2023

Towards Authentic Face Restoration with Iterative Diffusion Models and Beyond

Yang Zhao, Tingbo Hou, Yu-Chuan Su +2

An authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Mos…