2 papers
cs.LG2025
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
Feiyang Kang, Michael Kuchnik, Karthik Padthe +4
In post-training for reasoning Large Language Models (LLMs), the current state of practice trains LLMs in two independent stages: Supervised Fine-Tuning (SFT) and Reinforcement Lea…
cs.RO2025
Divide, Discover, Deploy: Factorized Skill Learning with Symmetry and Style Priors
Rafael Cathomen, Mayank Mittal, Marin Vlastelica +1
Unsupervised Skill Discovery (USD) allows agents to autonomously learn diverse behaviors without task-specific rewards. While recent USD methods have shown promise, their applicati…