3 papers
cs.LG2025
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Luca Baroni, Galvin Khara, Joachim Schaeffer +2
Layer-wise normalization (LN) is an essential component of virtually all transformer-based large language models. While its effects on training stability are well documented, its r…
cs.CV2025
Counterfactual contrastive learning: robust representations via causal image synthesis
Melanie Roschewitz, Fabio De Sousa Ribeiro, Tian Xia +2
Contrastive pretraining is well-known to improve downstream task performance and model generalisation, especially in limited label settings. However, it is sensitive to the choice…
cs.CV2025
Robust image representations with counterfactual contrastive learning
Mélanie Roschewitz, Fabio De Sousa Ribeiro, Tian Xia +2
Contrastive pretraining can substantially increase model generalisation and downstream performance. However, the quality of the learned representations is highly dependent on the d…