activity
20242026
collaborators

23 papers

cs.LG2026

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov

It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particul…

cs.AI2026

VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models

Damir Shodiev, Aleksei Staroverov, Nikita Kachaev +2

Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often…

cs.LG2026

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev +1

Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle mu…

cs.AI2026

Imagine to Ensure Safety in Hierarchical Reinforcement Learning

Gregory Gorbov, Artem Latyshev, Aleksandr I. Panov

This work investigates the safe exploration problem in reinforcement learning, where an agent must maximize cumulative performance while simultaneously satisfying safety constraint…

cs.LG2026

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10

Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual kno…

cs.LG2026

VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

Egor Cherepanov, Nikita Kachaev, Daniil Zelezetsky +6

Vision-language-action (VLA) models predict chunks of future actions from the current observation, an assumption that fails under partial observability, where decisions depend on i…