4 papers
NeST: Neuron Selective Tuning for LLM Safety
Sasha Behrouzi, Lichao Wu, Mohamadreza Rostami +1
Safety alignment is essential for the responsible deployment of Large Language Models (LLMs). Yet, existing approaches often rely on heavyweight fine-tuning that is costly to updat…
GoodVibe: Security-by-Vibe for LLM-Based Code Generation
Maximilian Thang, Lichao Wu, Sasha Behrouzi +4
Large language models (LLMs) are increasingly used for code generation in fast, informal development workflows, often referred to as vibe coding, where speed and convenience are pr…
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
Lichao Wu, Sasha Behrouzi, Mohamadreza Rostami +3
Safety alignment is critical for the ethical deployment of large language models (LLMs), guiding them to avoid generating harmful or unethical content. Current alignment techniques…
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
Lichao Wu, Sasha Behrouzi, Mohamadreza Rostami +2
Mixture-of-Experts (MoE) architectures have advanced the scaling of Large Language Models (LLMs) by activating only a sparse subset of parameters per input, enabling state-of-the-a…