11 papers
ERank in Latent Space as an Image-Complexity and Richness Measure
Maksim Smirnov, Grigory Kononov, Anastasiia Linich +2
We propose the effective rank (ERank) of the channel covariance of an image's deep feature map as a per-sample, label-free measure of visual richness, computed from a single forwar…
Geometric Metrics and LLMs: What They Measure and When They Work
Viacheslav Yusupov, Anna Antipina, Ameliia Alaeva +6
We present a systematic stress-test of geometric metrics for LLM evaluation. Rank-based geometric properties of internal representations have shown promise as reference-free qualit…
Bug or Feature: Weight Drift, Activation Sparsity and Spikes
Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav +3
The design of modern neural architectures has converged through incremental empirical choices, yet the mechanisms governing their training dynamics remain only partially understood…
Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
Maxim Zhelnin, Dmitry Redko, Daniil Volkov +8
Sequential recommendations (SR) with transformer-based architectures are widely adopted in real-world applications, where SR models require frequent retraining to adapt to ever-cha…
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
Shirin Alanova, Kristina Kazistova, Ekaterina Galaeva +7
The demand for efficient large language model (LLM) inference has intensified the focus on sparsification techniques. While semi-structured (N:M) pruning is well-established for we…
From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
Egor Maximov, Yulia Kuzkina, Azamat Kanametov +4
As large language models (LLMs) grow in size, efficient compression techniques like quantization and sparsification are critical. While quantization maintains performance with redu…