Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Bug or Feature: Weight Drift, Activation Sparsity and Spikes
Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav +3
The design of modern neural architectures has converged through incremental empirical choices, yet the mechanisms governing their training dynamics remain only partially understood…
cs.LG2026
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
Shirin Alanova, Kristina Kazistova, Ekaterina Galaeva +7
The demand for efficient large language model (LLM) inference has intensified the focus on sparsification techniques. While semi-structured (N:M) pruning is well-established for we…