6 papers
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
Congyi Fan, Jian Guan, Youtian Lin +5
Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction…
Dual-level Progressive Hardness-Aware Reweighting for Cross-View Geo-Localization
Guozheng Zheng, Jian Guan, Mingjie Xie +4
Cross-view geo-localization (CVGL) between drone and satellite imagery remains challenging due to severe viewpoint gaps and the presence of hard negatives, which are visually simil…
WaMaIR: Image Restoration via Multiscale Wavelet Convolutions and Mamba-based Channel Modeling with Texture Enhancement
Shengyu Zhu, Congyi Fan, Fuxuan Zhang
Image restoration is a fundamental and challenging task in computer vision, where CNN-based frameworks demonstrate significant computational efficiency. However, previous CNN-based…
DualMark: Identifying Model and Training Data Origins in Generated Audio
Xuefeng Yang, Jian Guan, Feiyang Xiao +5
Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to tr…
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
Congyi Fan, Jian Guan, Xuanjia Zhao +5
Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that natura…
FastDrag: Manipulate Anything in One Step
Xuanjia Zhao, Jian Guan, Congyi Fan +4
Drag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevail…