Publications (13)
SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
Viktoriia Zinkovich, Anton Antonov, Andrei Spiridonov +6
Multimodal large language models (MLLMs) have shown impressive capabilities in vision-language tasks such as reasoning segmentation, where models generate segmentation masks based…
BREPS: Bounding-Box Robustness Evaluation of Promptable Segmentation
Andrey Moskalenko, Danil Kuznetsov, Irina Dudko +6
Promptable segmentation models such as SAM have established a powerful paradigm, enabling strong generalization to unseen objects and domains with minimal user input, including poi…
TETRIS: Towards Exploring the Robustness of Interactive Segmentation
Andrey Moskalenko, Vlad Shakhuro, Anna Vorontsova +5
Interactive segmentation methods rely on user inputs to iteratively update the selection mask. A click specifying the object of interest is arguably the most simple and intuitive i…
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual kno…
Bridging the Gap Between Saliency Prediction and Image Quality Assessment
Kirillov Alexey, Andrey Moskalenko, Dmitriy Vatolin
Over the past few years, deep neural models have made considerable advances in image quality assessment (IQA). However, the underlying reasons for their success remain unclear, owi…
Deep Two-Stage High-Resolution Image Inpainting
Andrey Moskalenko, Mikhail Erofeev, Dmitriy Vatolin
In recent years, the field of image inpainting has developed rapidly, learning based approaches show impressive results in the task of filling missing parts in an image. But most d…
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
Andrey Moskalenko, Alexey Bryncev, Ivan Kosmynin +40
This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction me…
Bring the Apple, Not the Sofa: Impact of Irrelevant Context in Embodied AI Commands on VLA Models
Daria Pugacheva, Andrey Moskalenko, Denis Shepelev +3
Vision Language Action (VLA) models are widely used in Embodied AI, enabling robots to interpret and execute language instructions. However, their robustness to natural language va…
Angle-resolved time delay in photoemission
Jonas Wätzel, Andrey Moskalenko, Yaroslav Pavlyukh +1
We investigate theoretically the relative time delay of photoelectrons originating from different atomic subshells of noble gases. This quantity was measured via attosecond streaki…
AIM 2024 Challenge on Video Saliency Prediction: Methods and Results
Andrey Moskalenko, Alexey Bryncev, Dmitry Vatolin +30
This paper reviews the Challenge on Video Saliency Prediction at AIM 2024. The goal of the participants was to develop a method for predicting accurate saliency maps for the provid…
NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results
Nikolay Safonov, Alexey Bryncev, Andrey Moskalenko +28
This paper presents an overview of the NTIRE 2025 Challenge on UGC Video Enhancement. The challenge constructed a set of 150 user-generated content videos without reference ground…
RClicks: Realistic Click Simulation for Benchmarking Interactive Segmentation
Anton Antonov, Andrey Moskalenko, Denis Shepelev +4
The emergence of Segment Anything (SAM) sparked research interest in the field of interactive segmentation, especially in the context of image editing tasks and speeding up data an…
Temporally Coherent Person Matting Trained on Fake-Motion Dataset
Ivan Molodetskikh, Mikhail Erofeev, Andrey Moskalenko +1
We propose a novel neural-network-based method to perform matting of videos depicting people that does not require additional user input such as trimaps. Our architecture achieves…