activity
20242026
collaborators

6 papers

cs.LG2026

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

Leon Bergen, Usha Bhalla, Sidharth Baskaran +14

Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. Th…

cs.LG2026

Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training

Frank Xiao, Santiago Aranguri

We propose probe-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-differe…

cs.LG2026

Inference-Time Toxicity Mitigation in Protein Language Models

Manuel Fernández Burda, Santiago Aranguri, Iván Arcuschin Moreno +1

Protein language models (PLMs) are becoming practical tools for de novo protein design, yet their dual-use potential raises safety concerns. We show that domain adaptation to speci…

cs.LG2025

Phase-aware Training Schedule Simplifies Learning in Flow-Based Generative Models

Santiago Aranguri, Francesco Insulla

We analyze the training of a two-layer autoencoder used to parameterize a flow-based generative model for sampling from a high-dimensional Gaussian mixture. Previous work shows tha…

cs.LG2025

Optimizing Noise Schedules of Generative Models in High Dimensionss

Santiago Aranguri, Giulio Biroli, Marc Mezard +1

Recent works have shown that diffusion models can undergo phase transitions, the resolution of which is needed for accurately generating samples. This has motivated the use of diff…

cs.LG2024

Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes

Zhenfeng Tu, Santiago Aranguri, Arthur Jacot

The training dynamics of linear networks are well studied in two distinct setups: the lazy regime and balanced/active regime, depending on the initialization and width of the netwo…