1 citations · 1 across the 4 of their papers we have counts for
4 papers
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi +9
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly cond…
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti +5
Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and subsequently apply fine-grain…
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
Federico Alvetreti, Jary Pomponi, Paolo Di Lorenzo +1
This paper proposes a novel communication-efficient Split Learning (SL) framework, named Attention-based Double Compression (ADC), which reduces the communication overhead required…
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
Alessio Devoto, Federico Alvetreti, Jary Pomponi +3
Recently, foundation models based on Vision Transformers (ViTs) have become widely available. However, their fine-tuning process is highly resource-intensive, and it hinders their…