4 papers
Reconstruction-Shift Discrimination via Mask-Guided Latent Diffusion for Medical Anomaly Detection
Yibo Wan, Jinyu Cai, Yunhe Zhang +2
Unsupervised medical anomaly detection learns normal anatomical patterns from healthy training images and identifies deviations at test time. Reconstruction-based and diffusion-bas…
CFALR: Collaborative Filtering-Augmented Large Language Model for Personalized Fashion Outfit Recommendation
Yujuan Ding, Junrong Liao, Yunshan Ma +4
Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibil…
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
Yifan Xie, Fei Ma, Yi Bin +2
Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the signific…
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
Haoxuan Li, Yi Bin, Yunshan Ma +4
Cross-modal retrieval (CMR) is a fundamental task in multimedia research, focused on retrieving semantically relevant targets across different modalities. While traditional CMR met…