9 papers
SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation
Ruijie Sang, Yiqun Duan, Pinhan Fu +3
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…
Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation
Na Sang, Ding Ma, Rui Sang +1
Few-shot prompt learning is an effective strategy for adapting CLIP to downstream tasks, but class-only prompt optimization can overfit base-class supervision and weaken transfer t…
Membership Inference Attack Against Music Diffusion Models via Generative Manifold Perturbation
Yuxuan Liu, Peihong Zhang, Rui Sang +4
Membership inference attacks (MIAs) test whether a specific audio clip was used to train a model, making them a key tool for auditing generative music models for copyright complian…
TopSeg: A Multi-Scale Topological Framework for Data-Efficient Heart Sound Segmentation
Peihong Zhang, Zhixin Li, Yuxuan Liu +4
Deep learning approaches for heart-sound (PCG) segmentation built on time-frequency features can be accurate but often rely on large expert-labeled datasets, limiting robustness an…
DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification under Domain Shift
Peihong Zhang, Yuxuan Liu, Rui Sang +4
Acoustic scene classification (ASC) suffers from device-induced domain shift, especially when labels are limited. Prior work focuses on curriculum-based training schedules that str…
SceneGuard: Training-Time Voice Protection with Scene-Consistent Audible Background Noise
Rui Sang, Yuxuan Liu
Voice cloning technology poses significant privacy threats by enabling unauthorized speech synthesis from limited audio samples. Existing defenses based on imperceptible adversaria…