activity
20242026
most citedMixture-of-Experts Graph Transformers for Interpretable Particle Collision Detection

2 citations · 2 across the 1 of their papers we have counts for

collaborators

7 papers

cs.LG20262 cited

Mixture-of-Experts Graph Transformers for Interpretable Particle Collision Detection

Donatella Genovese, Alessandro Sgroi, Alessio Devoto +6

The Large Hadron Collider at CERN produces immense volumes of complex data from high-energy particle collisions, demanding sophisticated analytical techniques for effective interpr…

cs.LG2026

Universal Properties of Activation Sparsity in Modern Large Language Models

Filip Szatkowski, Patryk Będkowski, Alessio Devoto +5

Activation sparsity is an intriguing property of deep neural networks that has been extensively studied in ReLU-based models, due to its advantages for efficiency, robustness, and…

cs.LG2025

Communication Efficient Split Learning of ViTs with Attention-based Double Compression

Federico Alvetreti, Jary Pomponi, Paolo Di Lorenzo +1

This paper proposes a novel communication-efficient Split Learning (SL) framework, named Attention-based Double Compression (ADC), which reduces the communication overhead required…

cs.CE2025

Interpretable Classification of Levantine Ceramic Thin Sections via Neural Networks

Sara Capriotti, Alessio Devoto, Simone Scardapane +2

Classification of ceramic thin sections is fundamental for understanding ancient pottery production techniques, provenance, and trade networks. Although effective, traditional petr…

cs.LG2025

Adaptive Semantic Token Communication for Transformer-based Edge Inference

Alessio Devoto, Jary Pomponi, Mattia Merluzzi +2

This paper presents an adaptive framework for edge inference based on a dynamically configurable transformer-powered deep joint source channel coding (DJSCC) architecture. Motivate…

cs.CL2025

Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression

Nathan Godey, Alessio Devoto, Yu Zhao +4

Autoregressive language models rely on a Key-Value (KV) Cache, which avoids re-computing past hidden states during generation, making it faster. As model sizes and context lengths…