1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Kyungmin Park, Taesup Kim
Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shal…