Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
Sohir Maskey, Philipp Scholl, Jonas Knupp +2
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subse…
cs.AI2026
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
Jonas Knupp, Jan Hendrik Metzen, Jeremias Bohn +2
Depth-recurrence facilitates latent reasoning by sharing parameters across depths. However, prior work lacks combined FLOP-, parameter-, and memory-matched baselines, underutilizes…