4 papers · 1 filter
One Adapter Pair per Model: A Universal Activation Interface for Language Models
Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han
Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered f…
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
Su-Hyeon Kim, Hyundong Jin, Yejin Lee +1
While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior for model safety analysis. For exa…
Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
Su-Hyeon Kim, Yo-Sub Han
Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficult to compare or transfer ac…
How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs
Su-Hyeon Kim, Hyundong Jin, Yejin Lee +1
Large Reasoning Models (LRMs) achieve remarkable success through explicit thinking steps, yet the thinking steps introduce a novel risk by potentially amplifying unsafe behaviors.…