From the 1 of 6 linked papers with an AI index.
6 papers
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Mingjie Xie, Guangjun He, Dongli Xu +5
The paper introduces SynCLIP, a language‑image pretraining framework that aligns spatial attention across synonymous textual expressions to improve the robustness of open‑vocabular…
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
Congyi Fan, Jian Guan, Youtian Lin +5
Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction…
DualMark: Identifying Model and Training Data Origins in Generated Audio
Xuefeng Yang, Jian Guan, Feiyang Xiao +5
Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to tr…
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
Congyi Fan, Jian Guan, Xuanjia Zhao +5
Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that natura…
Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone Classification
Haiyan Lan, Shujun Li, Mingjie Xie +6
Local climate zone (LCZ) classification is of great value for understanding the complex interactions between urban development and local climate. Recent studies have increasingly f…
FastDrag: Manipulate Anything in One Step
Xuanjia Zhao, Jian Guan, Congyi Fan +4
Drag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevail…