20 papers
Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
Ruslan Rakhimov, George Bredis, Yuriy Maksyuta +1
Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at traini…
Rank-Then-Act: Reward-Free Control from Frame-Order Progress
Yuriy Maksyuta, George Bredis, Ruslan Rakhimov +1
We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a Vision-Language Model (VLM) o…
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky +3
Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training r…
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tok…
Trust-Region Behavior Blending for On-Policy Distillation
Daniil Plyusov, Alexey Gorbatovski, Alexey Malakhov +4
On-policy distillation (OPD) trains a student on prefixes sampled from its own policy while matching a stronger teacher. This addresses the prefix mismatch of offline distillation,…
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov +4
Reinforcement Learning with Verifiable Rewards (RLVR) is commonly based on group sampling to estimate advantages and stabilize policy updates. In practice, computational limits oft…