1 citations · 3 across the 10 of their papers we have counts for
6 papers · 1 filter
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
Grach Mkrtchian
First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without knowing whether a test clip comes from t…
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov +2
We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS syste…
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
Kirill Borodin, Vasiliy Kudryavtsev, Dmitrii Korzh +4
Automatic Speaker Verification (ASV) systems, which identify speakers based on their voice characteristics, have numerous applications, such as user authentication in financial tra…
Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative
Ksenia Lysikova, Kirill Borodin, Grach Mkrtchian
RuASD (Russian AntiSpoofing Dataset) is a dedicated, reproducible benchmark for Russian-language speech anti-spoofing designed to evaluate both in-domain discrimination and robustn…
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
Ivan Viakhirev, Kirill Borodin, Mikhail Gorodnichev +1
Multi-branch deep neural networks like AASIST3 achieve state-of-the-art comparable performance in audio anti-spoofing, yet their internal decision dynamics remain opaque compared t…
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov +3
We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by S…