7 papers
Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
Allison Zhuang, Santiago Aranguri
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is…
Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
Leon Bergen, Usha Bhalla, Sidharth Baskaran +14
Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. Th…
Inference-Time Toxicity Mitigation in Protein Language Models
Manuel Fernández Burda, Santiago Aranguri, Iván Arcuschin Moreno +1
Protein language models (PLMs) are becoming practical tools for de novo protein design, yet their dual-use potential raises safety concerns. We show that domain adaptation to speci…
Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
Frank Xiao, Santiago Aranguri
We propose probe-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-differe…
Optimizing Noise Schedules of Generative Models in High Dimensionss
Santiago Aranguri, Giulio Biroli, Marc Mezard +1
Recent works have shown that diffusion models can undergo phase transitions, the resolution of which is needed for accurately generating samples. This has motivated the use of diff…
Phase-aware Training Schedule Simplifies Learning in Flow-Based Generative Models
Santiago Aranguri, Francesco Insulla
We analyze the training of a two-layer autoencoder used to parameterize a flow-based generative model for sampling from a high-dimensional Gaussian mixture. Previous work shows tha…