4 papers
Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech
Vadim Popov, Wenju Gu, Tasnima Sadekova +2
Continuous diffusion for categorical data is a framework belonging to the diffusion family and aiming at generating discrete data. The scientific interest to such models has been c…
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
Georgii Aparin, Vadim Popov, Tasnima Sadekova +1
Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input. We invest…
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
Georgii Aparin, Tasnima Sadekova, Alexey Rukhovich +5
Sparse Autoencoders (SAEs) are powerful tools for interpreting neural representations, yet their use in audio remains underexplored. We train SAEs across all encoder layers of Whis…
Training-Free Voice Conversion with Factorized Optimal Transport
Alexander Lobashev, Assel Yermekova, Maria Larchenko
This paper introduces Factorized MKL-VC, a training-free modification for kNN-VC pipeline. In contrast with original pipeline, our algorithm performs high quality any-to-any cross-…