3 papers
cs.CV2026
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Team Kandinsky, Julia Agafonova, Bulat Akhmatov +85
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kan…
cs.SD2026
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko +4
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that…
cs.CV2026
KVAE: Family of Tokenizers for Multimodal Generative Models
Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part…