3 papers
cs.SD2026
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
Yonggang Zhu, Liting Gao, Aidong Men +1
Contrastive Language-Audio Pretraining (CLAP) models are widely used for audio understanding and support modality-agnostic condition swapping in many zero-shot applications. Howeve…
cs.CV2024
Filter or Compensate: Towards Invariant Representation from Distribution Shift for Anomaly Detection
Zining Chen, Xingshuang Luo, Weiqiu Wang +3
Recent Anomaly Detection (AD) methods have achieved great success with In-Distribution (ID) data. However, real-world data often exhibits distribution shift, causing huge performan…
cs.CV2024
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
Yuhang Zhang, Yuan Zhou, Zeyu Liu +4
Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limita…