3 papers
cs.CV2026
SSR: A Training-Free Approach for Streaming 3D Reconstruction
Hui Deng, Yuxin Mao, Yuxin He +1
Streaming 3D reconstruction demands long-horizon state updates under strict latency constraints, yet stateful recurrent models often suffer from geometric drift as errors accumulat…
cs.CV2025
Learning Spatial Decay for Vision Transformers
Yuxin Mao, Zhen Qin, Jinxing Zhou +4
Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spa…
cs.CV2025
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang +4
We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio a…