8 papers
NABLA: Neighborhood Adaptive Block-Level Attention
Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6
Recent progress in transformer-based architectures has demonstrated remarkable success in video generation tasks. However, the quadratic complexity of full attention mechanisms rem…
Making Time Editable in Video Diffusion Transformers
Konstantin Kuklev, Viacheslav Vasilev, Alexander Kunitsyn +2
Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics. We propose a temporal-control methodol…
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22
This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core lin…
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
Lev Novitskiy, Viacheslav Vasilev, Maria Kovaleva +2
Variational Autoencoders (VAEs) remain a cornerstone of generative computer vision, yet their training is often plagued by artifacts that degrade reconstruction and generation qual…
GHOST 2.0: generative high-fidelity one shot transfer of heads
Alexander Groshev, Anastasiia Iashchenko, Pavel Paramonov +2
While the task of face swapping has recently gained attention in the research community, a related problem of head swapping remains largely unexplored. In addition to skin color tr…
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
Viacheslav Vasilev, Julia Agafonova, Nikolai Gerasimenko +4
Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring…