32 papers
Understanding Layer Patching in Model Size Interpolation
Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak +3
Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distill…
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
Nicolas Anguita, Francesco Locatello, Andrew M. Saxe +4
Pretraining and fine-tuning are central stages in modern machine learning systems. In practice, feature learning plays an important role across both stages: deep neural networks le…
From Tokens to Policy: Causal and Interpretable Heterogeneous Treatment Effects Identification
Riccardo Cadei, Frank Otchere, Nyasha Tirivayi +3
Heterogeneous Treatment Effect (HTE) identification is crucial to explain the impact of an intervention and optimize our policies accordingly. Existing approaches trade expressivit…
Assessing Sample Quality in Conditional Generation under Compositional Shift
Berker Demirel, Valentino Maiorca, Marco Fumero +2
Conditional generators provide a natural tool for controllable generation, including settings where the desired condition is a new composition of observed attributes or experimenta…
The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench
Dingling Yao, Andrea Polesello, Adeel Pervez +2
While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invaria…
Toward Identifiable Sparse Autoencoders
Walter Nelson, Theofanis Karaletsos, Francesco Locatello
Recently, sparse autoencoders (SAEs) have emerged as an attractive tool for interpreting and interacting with representations in practical neural networks. While it is common empir…