6 citations · 11 across the 10 of their papers we have counts for
10 papers · 1 filter
Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability
Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi +1
Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previous works examined the effects…
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
Kathleen C. Fraser, Hillary Dawkins, Isar Nejadgholi +1
Fine-tuning a general-purpose large language model (LLM) for a specific domain or task has become a routine procedure for ordinary users. However, fine-tuning is known to remove th…
Gender-Neutral Machine Translation Strategies in Practice
Hillary Dawkins, Isar Nejadgholi, Chi-kiu Lo
Gender-inclusive machine translation (MT) should preserve gender ambiguity in the source to avoid misgendering and representational harms. While gender ambiguity often occurs natur…
When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text
Hillary Dawkins, Kathleen C. Fraser, Svetlana Kiritchenko
Detecting AI-generated text is a difficult problem to begin with; detecting AI-generated text on social media is made even more difficult due to the short text length and informal,…
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
Hillary Dawkins, Isar Nejadgholi, Chi-kiu Lo
We assess the difficulty of gender resolution in literary-style dialogue settings and the influence of gender stereotypes. Instances of the test suite contain spoken dialogue inter…
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
Rongchen Guo, Isar Nejadgholi, Hillary Dawkins +2
This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which dem…