2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023
Improving Activation Steering in Language Models with Mean-Centring
Ole Jorgensen, Dylan Cope, Nandi Schoots +1
Recent work in activation steering has demonstrated the potential to better control the outputs of Large Language Models (LLMs), but it involves finding steering vectors. This is d…
cs.CL2023★ 2 cited
Self-Consistency of Large Language Models under Ambiguity
Henning Bartsch, Ole Jorgensen, Domenic Rosati +2
Large language models (LLMs) that do not give consistent answers across contexts are problematic when used for tasks with expectations of consistency, e.g., question-answering, exp…