2 citations · 2 across the 5 of their papers we have counts for
7 papers · 1 filter
Making Time Editable in Video Diffusion Transformers
Konstantin Kuklev, Viacheslav Vasilev, Alexander Kunitsyn +2
Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics. We propose a temporal-control methodol…
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko +22
This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core lin…
NABLA: Neighborhood Adaptive Block-Level Attention
Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6
Recent progress in transformer-based architectures has demonstrated remarkable success in video generation tasks. However, the quadratic complexity of full attention mechanisms rem…
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
Lev Novitskiy, Viacheslav Vasilev, Maria Kovaleva +2
Variational Autoencoders (VAEs) remain a cornerstone of generative computer vision, yet their training is often plagued by artifacts that degrade reconstruction and generation qual…
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
Viacheslav Vasilev, Julia Agafonova, Nikolai Gerasimenko +4
Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring…
GHOST 2.0: generative high-fidelity one shot transfer of heads
Alexander Groshev, Anastasiia Iashchenko, Pavel Paramonov +2
While the task of face swapping has recently gained attention in the research community, a related problem of head swapping remains largely unexplored. In addition to skin color tr…