10 papers
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Maximo Rulli, Maximo Eduardo Rulli, Thomas Vaitses Fontanari +11
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly cond…
Steering Vectors are an Adversarial Attack Surface
Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3
Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning. Since the technique is plug-and-play, users share datasets and prec…
How Neural Losses Shape VAE Latents
Giorgio Strano, Luca Cerovaz, Michele Mancusi +2
Modern VAEs are rarely trained with the pointwise likelihood implied by the standard -VAE objective. In practice, pointwise reconstruction is often combined with perceptual and…
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
Tommaso Mencattini, Francesco Montagna, Francesco Locatello
Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the rate) often fail to learn mon…
Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
Mario Iacobelli, Adrian Robert Minut, Tommaso Mencattini +5
Reasoning models achieve strong performance on complex problems by leveraging long chains of thought, but this deliberate reasoning incurs substantial inference-time cost. The Long…
Language Models are Injective and Hence Invertible
Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi +3
Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could map to the same output and prevent exac…