3 papers
cs.LG2025
Universal Properties of Activation Sparsity in Modern Large Language Models
Filip Szatkowski, Patryk Będkowski, Alessio Devoto +5
Activation sparsity is an intriguing property of deep neural networks that has been extensively studied in ReLU-based models, due to its advantages for efficiency, robustness, and…
cs.LG2025
Rethinking Calibration for Early-Exit Neural Networks
Piotr Kubaty, Filip Szatkowski, Grzegorz Choczyński +2
Early-exit neural networks (EENNs) accelerate inference by allowing intermediate classifiers to stop computation once predictions are confident enough. Most methods rely on confide…
cs.LG2025
Efficient Multi-Source Knowledge Transfer by Model Merging
Marcin Osial, Bartosz Wójcik, Bartosz Zieliński +1
While transfer learning is an effective strategy, it often overlooks the opportunity to leverage knowledge from numerous available models online. Addressing this multi-source trans…