3 papers
cs.CR2026
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
Alex Polyakov, Daniel Kuznetsov
Safety alignment in large language models relies on behavioral training that can be overridden when sufficiently strong in-context patterns compete with learned refusal behaviors.…
cs.LG2026
Fairness-Aware Federated Learning with Trajectory Shapley Value
Daniel Kuznetsov, Ziqi Wang
Federated learning is an emerging distributed paradigm that addresses the challenges posed by heterogeneous, privacy-sensitive data. It enables multiple clients to train a model co…
cs.CR2026
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
Daniel Kuznetsov, Ofir Cohen, Karin Shistik +2
Safety-aligned LLMs go through refusal training to reject harmful requests, but whether these mechanisms remain effective under emotionally charged stimuli is unexplored. We introd…