1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2023
Quantization Aware Factorization for Deep Neural Network Compression
Daria Cherniuk, Stanislav Abukhovich, Anh-Huy Phan +3
Tensor decomposition of convolutional and fully-connected layers is an effective way to reduce parameters and FLOP in neural networks. Due to memory and power consumption limitatio…
cs.AI2023★ 1 cited
Efficient GPT Model Pre-training using Tensor Train Matrix Representation
Viktoriia Chekalina, Georgii Novikov, Julia Gusak +2
Large-scale transformer models have shown remarkable performance in language modelling tasks. However, such models feature billions of parameters, leading to difficulties in their…