Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@ Crossover on a Free-Verifier Domain
Igor Lima Strozzi
When a language model trains on its own verified outputs, does it acquire capability beyond its base, or merely get better at expressing capability the base already had? We make th…
cs.LG2026
Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
Igor Strozzi
We measure procedural-skill SFT contribution across three Qwen3.5 dense scales (0.8B, 2B, 4B) on a 200-task / 40-skill holdout, with Claude Haiku 4.5 as a frontier reference. The c…