4 papers
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe +1
Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making rel…
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
Connor Baumler, Calvin Bao, Huy Nghiem +3
Despite the growing use of large language models (LLMs) for writing tasks, users may hesitate to rely on LLMs when personal style is important. Post-editing LLM-generated drafts or…
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
Dayeon Ki, Yu Hou, Rachel Rudinger +3
Neologisms and emerging slang are central to daily conversation, yet challenging for non-native speakers (NNS) to interpret and use appropriately in cross-cultural communication wi…
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
Lingjun Zhao, Dayeon Ki, Marine Carpuat +1
Language models are known to exhibit various forms of cultural bias in decision-making tasks, yet much less is known about their degree of cultural familiarity in open-ended text g…