2 papers
cs.CV2025
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang +4
We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio a…
cs.CV2024
Deep Non-rigid Structure-from-Motion: A Sequence-to-Sequence Translation Perspective
Hui Deng, Tong Zhang, Yuchao Dai +3
Directly regressing the non-rigid shape and camera pose from the individual 2D frame is ill-suited to the Non-Rigid Structure-from-Motion (NRSfM) problem. This frame-by-frame 3D re…