2 papers
cs.SD2026
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Tianyu Xu, Sieun Kim, Qianhui Zheng +6
In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce…
cs.CV2025
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
Beijia Lu, Ziyi Chen, Jing Xiao +1
Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods…