3 papers
cs.CV2026
KVAE: Family of Tokenizers for Multimodal Generative Models
Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part…
cs.SD2024
FINALLY: fast and universal speech enhancement with studio-like quality
Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev +6
In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, an…
eess.AS2024
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
Hanbin Bae, Pavel Andreev, Azat Saginbaev +4
This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations…