5 papers · 1 filter
LaPrune: Controllable Differentiable Sparsity at Million Scale
Jakub Antczak, Joanna Wojciechowicz, Łukasz Struski +1
Top- selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selec…
SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs
Mikołaj Zasada, Łukasz Struski, Jacek Tabor +1
Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subset of experts via top- routing. While t…
LAPLEX: The FFT of Learnable Laplace Kernels
Łukasz Struski, Hanna Blazhko, Piotr Kubaty +1
Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geometry paid for by dense paramete…
FeNeC: Enhancing Continual Learning via Feature Clustering with Neighbor- or Logit-Based Classification
Kamil Książek, Hubert Jastrzębski, Bartosz Trojan +3
The ability of deep learning models to learn continuously is essential for adapting to new data categories and evolving data distributions. In recent years, approaches leveraging f…
SEMU: Singular Value Decomposition for Efficient Machine Unlearning
Marcin Sendera, Łukasz Struski, Kamil Książek +3
While the capabilities of generative foundational models have advanced rapidly in recent years, methods to prevent harmful and unsafe behaviors remain underdeveloped. Among the pre…