16 papers
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle +10
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, d…
Asking Back: Interaction-Layer Antidistillation Watermarks
Guang Yang, Amir Ghasemian, Fengchen Liu +3
Detecting unauthorized knowledge distillation from a deployed LLM API is hard because the defender controls neither the attacker's training pipeline nor the next-token logits. Exis…
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
Asymmetric Phase Coding Audio Watermarking
Guang Yang, Amir Ghasemian, Ninareh Mehrabi +1
The proliferation of deepfake audio challenges voice-based authentication systems; passive forensic detectors are sensitive to evolving generative models and to real-world channel…
SWAN: Semantic Watermarking with Abstract Meaning Representation
Ziping Ye, Gourab Dey, Christos Christodoulopoulos +7
We introduce SWAN (Semantic Watermarking with Abstract Meaning Representation), a novel framework that embeds watermark signatures into the semantic structure of a sentence using A…
FERRET: Framework for Expansion Reliant Red Teaming
Ninareh Mehrabi, Vitor Albiero, Maya Pavlova +1
We introduce a multi-faceted automated red teaming framework in which the goal is to generate multi-modal adversarial conversations that would break a target model and introduce va…