2 papers
cs.LG2026
Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
Roman Maksimov, Vladimir Aletov, Vladimir Solodkin +3
As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspire…
cs.LG2026
Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates
Roman Maksimov, Vladimir Aletov, Dmitry Bylinkin +3
Knowledge editing (KE) provides a lightweight alternative to repeated fine-tuning of LLMs. However, most existing KE methods target dense feed-forward layers, while modern LLMs inc…