activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Bug or Feature: Weight Drift, Activation Sparsity and Spikes

Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav +3

The design of modern neural architectures has converged through incremental empirical choices, yet the mechanisms governing their training dynamics remain only partially understood…

cs.LG2026

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches

Shirin Alanova, Kristina Kazistova, Ekaterina Galaeva +7

The demand for efficient large language model (LLM) inference has intensified the focus on sparsification techniques. While semi-structured (N:M) pruning is well-established for we…

cs.LG2026

From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction

Egor Maximov, Yulia Kuzkina, Azamat Kanametov +4

As large language models (LLMs) grow in size, efficient compression techniques like quantization and sparsification are critical. While quantization maintains performance with redu…

cs.LG2025

How to model Human Actions distribution with Event Sequence Data

Egor Surkov, Dmitry Osin, Evgeny Burnaev +1

This paper studies forecasting of the future distribution of events in human action sequences, a task essential in domains like retail, finance, healthcare, and recommendation syst…

cs.LG2025

EBES: Easy Benchmarking for Event Sequences

Dmitry Osin, Igor Udovichenko, Viktor Moskvoretskii +2

Event Sequences (EvS) refer to sequential data characterized by irregular sampling intervals and a mix of categorical and numerical features. Accurate classification of these seque…

cs.LG2024

GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs

Maxim Zhelnin, Viktor Moskvoretskii, Egor Shvetsov +4

Parameter Efficient Fine-Tuning (PEFT) methods have gained popularity and democratized the usage of Large Language Models (LLMs). Recent studies have shown that a small subset of w…