5 papers
Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
Srikanth Malla, Chiho Choi, Joon Hee Choi
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that directio…
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
Srikanth Malla, Chiho Choi, Joon Hee Choi
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-s…
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
Sangmin Yoo, Srikanth Malla, Chiho Choi +2
The inference of large language models imposes significant computational workloads, often requiring the processing of billions of parameters. Although early-exit strategies have pr…
CAMILA: Context-Aware Masking for Image Editing with Language Alignment
Hyunseung Kim, Chiho Choi, Srikanth Malla +3
Text-guided image editing has been allowing users to transform and synthesize images through natural language instructions, offering considerable flexibility. However, most existin…
COPAL: Continual Pruning in Large Language Generative Models
Srikanth Malla, Joon Hee Choi, Chiho Choi
Adapting pre-trained large language models to different domains in natural language processing requires two key considerations: high computational demands and model's inability to…