Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Evaluating Language Models for Harmful Manipulation
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9
Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…
cs.AI2024
The effect of fine-tuning on language model toxicity
Will Hawkins, Brent Mittelstadt, Chris Russell
Fine-tuning language models has become increasingly popular following the proliferation of open models and improvements in cost-effective parameter efficient fine-tuning. However,…