3 citations · 3 across the 3 of their papers we have counts for
4 papers
PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models
Ruolan Sun, Pawel Polak
Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, attention editing, or reward-based…
Every Image Listens, Every Image Dances: Music-Driven Image Animation
Zhikang Dong, Weituo Hao, Ju-Chiang Wang +2
Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video g…
Face-GPS: A Comprehensive Technique for Quantifying Facial Muscle Dynamics in Videos
Juni Kim, Zhikang Dong, Pawel Polak
We introduce a novel method that combines differential geometry, kernels smoothing, and spectral analysis to quantify facial muscle activity from widely accessible video recordings…
MuseChat: A Conversational Music Recommendation System for Videos
Zhikang Dong, Bin Chen, Xiulong Liu +2
Music recommendation for videos attracts growing interest in multi-modal research. However, existing systems focus primarily on content compatibility, often ignoring the users' pre…