2 papers
cs.LG2026
User Model Extraction via Belief Self-Distillation
Ali Holmov, Yiran Huang, Kirill Bykov +1
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate.…
cs.LG2026
One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
Ali Holmov, Paul Youssef, Nandi Schoots +1
Knowledge editing methods such as ROME and MEMIT update factual associations in transformer models by modifying MLP weights. While evaluated mainly by output behavior, their intern…