3 papers
cs.CV2026
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Team Kandinsky, Julia Agafonova, Bulat Akhmatov +85
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kan…
cs.CV2026
VIBE: Visual Instruction Based Editor
Grigorii Alekseenko, Aleksandr Gordeev, Irina Tolstykh +7
Instruction-based image editing is among the fastest developing areas in generative AI. Over the past year, the field has reached a new level, with dozens of open-source models rel…
cs.CV2025
NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
Maksim Kuprashevich, Grigorii Alekseenko, Irina Tolstykh +4
Recent advances in generative modeling enable image editing assistants that follow natural language instructions without additional user input. Their supervised training requires m…