collaborators

5 papers

cs.LG2026

Vision-Language Models Unlock Task-Centric Latent Actions

Alexander Nikulin, Ilya Zisman, Albina Klepach +5

Latent Action Models (LAMs) have rapidly gained traction as an important component in the pre-training pipelines of leading Vision-Language-Action models. However, they fail when o…

cs.CV2025

NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows

Denis Tarasov, Alexander Nikulin, Ilya Zisman +5

Recent advances in Vision-Language-Action (VLA) models have established a two-component architecture, where a pre-trained Vision-Language Model (VLM) encodes visual observations an…

cs.CV2025

Object-Centric Latent Action Learning

Albina Klepach, Alexander Nikulin, Ilya Zisman +6

Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distrac…

cs.CV2025

Latent Action Learning Requires Supervision in the Presence of Distractors

Alexander Nikulin, Ilya Zisman, Denis Tarasov +4

Recently, latent action learning, pioneered by Latent Action Policies (LAPO), have shown remarkable pre-training efficiency on observation-only data, offering potential for leverag…

cs.LG2025

Vintix: Action Model via In-Context Reinforcement Learning

Andrey Polubarov, Nikita Lyubaykin, Alexander Derevyagin +4

In-Context Reinforcement Learning (ICRL) represents a promising paradigm for developing generalist agents that learn at inference time through trial-and-error interactions, analogo…