4 papers
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko +4
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that…
KVAE: Family of Tokenizers for Multimodal Generative Models
Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part…
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22
This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core lin…
Hierarchical B-frame Video Coding for Long Group of Pictures
Ivan Kirillov, Denis Parkhomenko, Kirill Chernyshev +4
Learned video compression methods already outperform VVC in the low-delay (LD) case, but the random-access (RA) scenario remains challenging. Most works on learned RA video compres…