5 papers · 1 filter
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur +6
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing…
Beyond Individual Personas: Aligning Synthetic Dialogue to Population-Level Behavior Distributions
Xinyi Liu, Rinat Khaziev, Hooshang Nayyeri +3
Synthetic dialogue corpora are increasingly used as proxies for target dialogue data, yet persona-grounded generators optimize individual conversations rather than corpus compositi…
SWAN: Semantic Watermarking with Abstract Meaning Representation
Ziping Ye, Gourab Dey, Christos Christodoulopoulos +7
We introduce SWAN (Semantic Watermarking with Abstract Meaning Representation), a novel framework that embeds watermark signatures into the semantic structure of a sentence using A…
Evaluating Differentially Private Synthetic Data Generation in High-Stakes Domains
Krithika Ramesh, Nupoor Gandhi, Pulkit Madaan +3
The difficulty of anonymizing text data hinders the development and deployment of NLP in high-stakes domains that involve private data, such as healthcare and social services. Poor…
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
Tao Meng, Ninareh Mehrabi, Palash Goyal +6
We propose a constraint learning schema for fine-tuning Large Language Models (LLMs) with attribute control. Given a training corpus and control criteria formulated as a sequence-l…