Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Diff Mining: Logit Differences Reveal Finetuning Objectives
Greg Kocher, Robert West, Clément Dumas +1
Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains unclear exactly which behaviors emerge during…
cs.LG2026
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder, Viktor Moskvoretskii, Raghav Singhal +12
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant…