1 citations · 1 across the 5 of their papers we have counts for
5 papers · 1 filter
VIBE: Visual Instruction Based Editor
Grigorii Alekseenko, Aleksandr Gordeev, Irina Tolstykh +7
Instruction-based image editing is among the fastest developing areas in generative AI. Over the past year, the field has reached a new level, with dozens of open-source models rel…
NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
Maksim Kuprashevich, Grigorii Alekseenko, Irina Tolstykh +4
Recent advances in generative modeling enable image editing assistants that follow natural language instructions without additional user input. Their supervised training requires m…
Saliency-Guided DETR for Moment Retrieval and Highlight Detection
Aleksandr Gordeev, Vladimir Dokholyan, Irina Tolstykh +1
Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited…
CerberusDet: Unified Multi-Dataset Object Detection
Irina Tolstykh, Mikhail Chernyshov, Maksim Kuprashevich
Conventional object detection models are usually limited by the data on which they were trained and by the category logic they define. With the recent rise of Language-Visual Model…
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
Maksim Kuprashevich, Grigorii Alekseenko, Irina Tolstykh
Multimodal Large Language Models (MLLMs) have recently gained immense popularity. Powerful commercial models like ChatGPT-4V and Gemini, as well as open-source ones such as LLaVA,…