2 papers
cs.LG2025
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
Liana Mikaelyan, Ayyoob Imani, Mathew Salvaris +2
We introduce DeltaLLM, a new post-training compression technique to reduce the memory footprint of LLMs. We propose an alternative way of structuring LLMs with weight sharing betwe…
cs.AI2024
KBLaM: Knowledge Base augmented Language Model
Xi Wang, Taketomo Isazawa, Liana Mikaelyan +1
In this paper, we propose Knowledge Base augmented Language Model (KBLaM), a new method for augmenting Large Language Models (LLMs) with external knowledge. KBLaM works with a know…