28 citations · 31 across the 8 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
Srikanth Malla, Chiho Choi, Joon Hee Choi
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that directio…
cs.LG2026
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
Srikanth Malla, Chiho Choi, Joon Hee Choi
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-s…
cs.LG2024
COPAL: Continual Pruning in Large Language Generative Models
Srikanth Malla, Joon Hee Choi, Chiho Choi
Adapting pre-trained large language models to different domains in natural language processing requires two key considerations: high computational demands and model's inability to…