4 papers · 1 filter
The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints
Yupei Li, Chenyang Lyu, Longyue Wang +4
Speech deepfake detection (DFD) has benefited from diverse acoustic and semantic speech representations, many of which encode valuable speech information and are costly to train. P…
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
Fei Yang, Xuanfan Ni, Renyi Yang +7
Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription,…
Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
Xuanfan Ni, Fei Yang, Fengping Tian +6
Automatic Speech Recognition (ASR) models have achieved remarkable accuracy in general settings, yet their performance often degrades in domain-specific applications due to data mi…
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
Haoqin Sun, Chenyang Lyu, Xiangyu Kong +9
Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discr…