works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.CV2026

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

Mingjie Xie, Guangjun He, Dongli Xu +5

The paper introduces SynCLIP, a language‑image pretraining framework that aligns spatial attention across synonymous textual expressions to improve the robustness of open‑vocabular…

cs.SD2026

Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling

Congyi Fan, Jian Guan, Youtian Lin +5

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction…

cs.SD2025

DualMark: Identifying Model and Training Data Origins in Generated Audio

Xuefeng Yang, Jian Guan, Feiyang Xiao +5

Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to tr…

cs.MM2025

Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation

Congyi Fan, Jian Guan, Xuanjia Zhao +5

Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that natura…

cs.CV2024

FastDrag: Manipulate Anything in One Step

Xuanjia Zhao, Jian Guan, Congyi Fan +4

Drag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevail…