4 papers
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22
This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core lin…
Unlocking the Duality between Flow and Field Matching
Daniil Shlenskii, Alexander Varlamov, Nazar Buzun +1
Conditional Flow Matching (CFM) unifies conventional generative paradigms such as diffusion models and flow matching. Interaction Field Matching (IFM) is a newer framework that gen…
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
Daria Diatlova, Nikita Balagansky, Alexander Varlamov +1
Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and stat…
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
Elvir Karimov, Alexander Varlamov, Danil Ivanov +2
Deep learning voice models are commonly used nowadays, but the safety processing of personal data, such as human identity and speech content, remains suspicious. To prevent malicio…