2 citations · 4 across the 3 of their papers we have counts for
5 papers · 1 filter
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22
This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core lin…
Kandinsky 3.0 Technical Report
Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6
We present Kandinsky 3.0, a large-scale text-to-image generation model based on latent diffusion, continuing the series of text-to-image Kandinsky models and reflecting our progres…
Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion
Anton Razzhigaev, Arseniy Shakhmatov, Anastasia Maltseva +7
Text-to-image generation is a significant domain in modern computer vision and has achieved substantial improvements through the evolution of generative architectures. Among these,…
RuCLIP -- new models and experiments: a technical report
Alex Shonenkov, Andrey Kuznetsov, Denis Dimitrov +10
In the report we propose six new implementations of ruCLIP model trained on our 240M pairs. The accuracy results are compared with original CLIP model with Ru-En translation (OPUS-…
A new face swap method for image and video domains: a technical report
Daniil Chesakov, Anastasia Maltseva, Alexander Groshev +2
Deep fake technology became a hot field of research in the last few years. Researchers investigate sophisticated Generative Adversarial Networks (GAN), autoencoders, and other appr…