1 paper · 1 filter
Joris Postmus, Steven Abreu
Large language models have transformed AI, yet reliably controlling their outputs remains a challenge. This paper explores activation engineering, where outputs of pre-trained LLMs…