7 papers
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe +1
Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making rel…
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
Sy-Tuyen Ho, Minghui Liu, Huy Nghiem +1
Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks ra…
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
Connor Baumler, Calvin Bao, Huy Nghiem +3
Despite the growing use of large language models (LLMs) for writing tasks, users may hesitate to rely on LLMs when personal style is important. Post-editing LLM-generated drafts or…
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
Huy Nghiem, Phuong-Anh Nguyen-Le, Sy-Tuyen Ho +1
Research has documented LLMs' name-based bias in hiring and salary recommendations. In this paper, we instead consider a setting where LLMs generate candidate summaries for downstr…
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
Huy Nghiem, Advik Sachdeva, Hal Daumé
WARNING: This paper contains examples of offensive materials. To address the proliferation of toxic content on social media, we introduce SMARTER, we introduce SMARTER, a data-effi…
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
Huy Nghiem, Swetasudha Panda, Devashish Khatwani +3
Large Language Models (LLMs) are increasingly used in healthcare, yet ensuring their safety and trustworthiness remains a barrier to deployment. Conversational medical assistants m…