11 papers
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4
We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but…
Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge
Masoumeh Chapariniya, Aref Farhadipour, Sarah Ebling +2
We present our system for the BLEMORE Challenge at FG 2026 on blended emotion recognition with relative salience prediction. Our approach combines six encoder families through late…
TidyVoice 2026 Challenge Evaluation Plan
Aref Farhadipour, Jan Marquenie, Srikanth Madikeri +6
The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To…
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
Aref Farhadipour, Teodora Vukovic, Volker Dellwo +2
Person identification systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently present with missing or degraded modalities. To address this cha…
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
Aref Farhadipour, Ming Jin, Valeriia Vyshnevetska +3
This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing at…
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
Aref Farhadipour, Jan Marquenie, Srikanth Madikeri +1
The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-sp…