collaborators

20 papers

cs.LG2026

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

Ruslan Rakhimov, George Bredis, Yuriy Maksyuta +1

Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at traini…

cs.LG2026

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

Yuriy Maksyuta, George Bredis, Ruslan Rakhimov +1

We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a Vision-Language Model (VLM) o…

cs.LG2026

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky +3

Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training r…

cs.LG2026

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1

Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tok…

cs.LG2026

Trust-Region Behavior Blending for On-Policy Distillation

Daniil Plyusov, Alexey Gorbatovski, Alexey Malakhov +4

On-policy distillation (OPD) trains a student on prefixes sampled from its own policy while matching a stronger teacher. This addresses the prefix mismatch of offline distillation,…

cs.LG2026

F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov +4

Reinforcement Learning with Verifiable Rewards (RLVR) is commonly based on group sampling to estimate advantages and stabilize policy updates. In practice, computational limits oft…