Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Weight-Space Geometry of Offline Reasoning Training
Aleksandr Nikolich, Igor Kiselev, Vladimir Platonov +1
Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared…
cs.LG2026
Yes, Q-learning Helps Offline In-Context RL
Denis Tarasov, Alexander Nikulin, Ilya Zisman +6
Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL set…