Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories
Kyungmin Park, Taesup Kim
Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shal…
cs.AI2026
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
Heejun Kim, Seungpil Lee, Jewon Yeom +5
Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from the learned content or from the act of inje…