4 papers · 1 filter
TensorSLM: Energy-efficient Embedding Compression of Sub-billion Parameter Language Models on Low-end Devices
Mingxue Xu, Yao Lei Xu, Danilo P. Mandic
Small Language Models (SLMs, or on-device LMs) have significantly fewer parameters than Large Language Models (LLMs). They are typically deployed on low-end devices, like mobile ph…
Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models
Harry J. Davies, Giorgos Iacovides, Danilo P. Mandic
The sheer scale of data required to train modern large language models (LLMs) poses significant risks, as models are likely to gain knowledge of sensitive topics such as bio-securi…
Geometry is All You Need: A Unified Taxonomy of Matrix and Tensor Factorization for Compression of Generative Language Models
Mingxue Xu, Sadia Sharmin, Danilo P. Mandic
Matrix and tensor-guided parametrization for Natural Language Processing (NLP) models is fundamentally useful for the improvement of the model's systematic efficiency. However, the…
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition
Mingxue Xu, Yao Lei Xu, Danilo P. Mandic
High-dimensional token embeddings underpin Large Language Models (LLMs), as they can capture subtle semantic information and significantly enhance the modelling of complex language…