Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
Lama Alssum, Hani Itani, Hasan Abed Al Kader Hammoud +3
The safety alignment of large language models (LLMs) is becoming increasingly important with their democratization. In this paper, we study the safety degradation that comes with a…
cs.CL2025
Shh, don't say that! Domain Certification in LLMs
Cornelius Emde, Alasdair Paren, Preetham Arvind +6
Large language models (LLMs) are often deployed to perform constrained tasks, with narrow domains. For example, customer support bots can be built on top of LLMs, relying on their…
cs.CL2024
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
Hasan Abed Al Kader Hammoud, Umberto Michieli, Fabio Pizzati +4
Merging Large Language Models (LLMs) is a cost-effective technique for combining multiple expert LLMs into a single versatile model, retaining the expertise of the original ones. H…