Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
EmbBERT: Attention Under 2 MB Memory
Riccardo Bravin, Massimo Pavan, Hazem Hesham Yousef Shalby +2
Transformer architectures based on the attention mechanism have revolutionized natural language processing (NLP), driving major breakthroughs across virtually every NLP task. Howev…
cs.CL2025
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
Miguel Nogales, Matteo Gambella, Manuel Roveri
Early exits (EEs) offer a promising approach to reducing computational costs and latency by dynamically terminating inference once a satisfactory prediction confidence on a data sa…