3 papers
cs.CV2025
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
Yingjie Xia, Xi Wang, Jinglei Shi +2
Images evoke emotions that profoundly influence perception, often prioritized over content. Current Image Emotional Synthesis (IES) approaches artificially separate generation and…
cs.SD2025
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
Ming Meng, Ziyi Yang, Jian Yang +3
Recent advancements in text-to-speech (TTS) technology have increased demand for personalized audio synthesis. Zero-shot voice cloning, a specialized TTS task, aims to synthesize a…
cs.CV2025
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
Junxian Ma, Shiwen Wang, Jian Yang +6
Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization. However, existing methods typically rely on constrained audio-visual align…