1 paper
Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee +3
Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge t…