3 papers
cs.CV2026
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Team Kandinsky, Julia Agafonova, Bulat Akhmatov +85
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kan…
cs.CV2026
Making Time Editable in Video Diffusion Transformers
Konstantin Kuklev, Viacheslav Vasilev, Alexander Kunitsyn +2
Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics. We propose a temporal-control methodol…
cs.CV2024
VLRM: Vision-Language Models act as Reward Models for Image Captioning
Maksim Dzabraev, Alexander Kunitsyn, Andrei Ivaniuta
In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and B…