6 papers
Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
Leon Bergen, Usha Bhalla, Sidharth Baskaran +14
Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. Th…
Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
Frank Xiao, Santiago Aranguri
We propose probe-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-differe…
Inference-Time Toxicity Mitigation in Protein Language Models
Manuel Fernández Burda, Santiago Aranguri, Iván Arcuschin Moreno +1
Protein language models (PLMs) are becoming practical tools for de novo protein design, yet their dual-use potential raises safety concerns. We show that domain adaptation to speci…
Phase-aware Training Schedule Simplifies Learning in Flow-Based Generative Models
Santiago Aranguri, Francesco Insulla
We analyze the training of a two-layer autoencoder used to parameterize a flow-based generative model for sampling from a high-dimensional Gaussian mixture. Previous work shows tha…
Optimizing Noise Schedules of Generative Models in High Dimensionss
Santiago Aranguri, Giulio Biroli, Marc Mezard +1
Recent works have shown that diffusion models can undergo phase transitions, the resolution of which is needed for accurately generating samples. This has motivated the use of diff…
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
Zhenfeng Tu, Santiago Aranguri, Arthur Jacot
The training dynamics of linear networks are well studied in two distinct setups: the lazy regime and balanced/active regime, depending on the initialization and width of the netwo…