4 papers
Weight-Space Geometry of Offline Reasoning Training
Aleksandr Nikolich, Igor Kiselev, Vladimir Platonov +1
Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared…
Yes, Q-learning Helps Offline In-Context RL
Denis Tarasov, Alexander Nikulin, Ilya Zisman +6
Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL set…
CayleyPy-4: AI-Holography. Towards analogs of holographic string dualities for AI tasks
A. Chervov, F. Levkovich-Maslyuk, A. Smolensky +41
This is the fourth paper in the CayleyPy project, which applies AI methods to the exploration of large graphs. In this work, we suggest the existence of a new discrete version of h…
Latent Action Learning Requires Supervision in the Presence of Distractors
Alexander Nikulin, Ilya Zisman, Denis Tarasov +4
Recently, latent action learning, pioneered by Latent Action Policies (LAPO), have shown remarkable pre-training efficiency on observation-only data, offering potential for leverag…