6 papers
SSR: A Training-Free Approach for Streaming 3D Reconstruction
Hui Deng, Yuxin Mao, Yuxin He +1
Streaming 3D reconstruction demands long-horizon state updates under strict latency constraints, yet stateful recurrent models often suffer from geometric drift as errors accumulat…
Learning Spatial Decay for Vision Transformers
Yuxin Mao, Zhen Qin, Jinxing Zhou +4
Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spa…
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang +4
We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio a…
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
Jinxing Zhou, Dan Guo, Yuxin Mao +3
Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identific…
You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
Zhen Qin, Yuxin Mao, Xuyang Shen +4
Linear attention mechanisms have gained prominence in causal language models due to their linear computational complexity and enhanced speed. However, the inherent decay mechanism…
TAVGBench: Benchmarking Text to Audible-Video Generation
Yuxin Mao, Xuyang Shen, Jing Zhang +5
The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both a…