5 papers
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
Jiaming Cheng, Ruiyu Liang, Ye Ni +6
In this paper, we propose an intra-set and inter-set recursive fusion framework with time-frequency calibrated knowledge distillation (ISRF-TFCKD) for SE. Different from previo…
Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation
Ye Ni, Ruiyu Liang, Xiaoshuai Hao +7
Hearing aids (HAs) are widely used to provide personalized speech enhancement (PSE) services, improving the quality of life for individuals with hearing loss. However, HA performan…
The First MPDD Challenge: Multimodal Personality-aware Depression Detection
Changzeng Fu, Zelin Fu, Qi Zhang +14
Depression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detectio…
Identity-free Artificial Emotional Intelligence via Micro-Gesture Understanding
Rong Gao, Xin Liu, Bohao Xing +3
In this work, we focus on a special group of human body language -- the micro-gesture (MG), which differs from the range of ordinary illustrative gestures in that they are not inte…
Gender Bias in Text-to-Video Generation Models: A case study of Sora
Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria +2
The advent of text-to-video generation models has revolutionized content creation as it produces high-quality videos from textual prompts. However, concerns regarding inherent bias…