From the 1 of 4 linked papers with an AI index.
4 papers
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Mingjie Xie, Guangjun He, Dongli Xu +5
The paper introduces SynCLIP, a language‑image pretraining framework that aligns spatial attention across synonymous textual expressions to improve the robustness of open‑vocabular…
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
Congyi Fan, Jian Guan, Youtian Lin +5
Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction…
DualMark: Identifying Model and Training Data Origins in Generated Audio
Xuefeng Yang, Jian Guan, Feiyang Xiao +5
Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to tr…
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
Congyi Fan, Jian Guan, Xuanjia Zhao +5
Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that natura…