7 papers
SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks
Adrian Robert Minut, Nico Daheim, Marco Miani +3
Structured weight-uncertainty can improve many aspects of deep learning, but it remains costly to estimate and difficult to implement. Here, we show that these issues can be addres…
Steering Vectors are an Adversarial Attack Surface
Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3
Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning. Since the technique is plug-and-play, users share datasets and prec…
Zero-Shot Quantization via Weight-Space Arithmetic
Daniele Solombrino, Antonio Andrea Gargiulo, Alessandro Zirilli +3
We show that robustness to post-training quantization (PTQ) is a transferable direction in weight space. We call this direction the quantization vector: extracted from a donor task…
Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
Mario Iacobelli, Adrian Robert Minut, Tommaso Mencattini +5
Reasoning models achieve strong performance on complex problems by leveraging long chains of thought, but this deliberate reasoning incurs substantial inference-time cost. The Long…
Spilled Energy in Large Language Models
Adrian Robert Minut, Hazem Dewidar, Iacopo Masi
We reinterpret the final Large Language Model (LLM) softmax classifier as an Energy-Based Model (EBM), decomposing the sequence-to-sequence probability chain into multiple interact…
Mergenetic: a Simple Evolutionary Model Merging Library
Adrian Robert Minut, Tommaso Mencattini, Andrea Santilli +2
Model merging allows combining the capabilities of existing models into a new one - post hoc, without additional training. This has made it increasingly popular thanks to its low c…