collaborators

11 papers

cs.CL2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur +5

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing…

cs.AI2026

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

Tharindu Kumarage, Lisa Bauer, Yao Ma +7

As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks w…

cs.CL2026

Beyond Individual Personas: Aligning Synthetic Dialogue to Population-Level Behavior Distributions

Xinyi Liu, Rinat Khaziev, Hooshang Nayyeri +3

Synthetic dialogue corpora are increasingly used as proxies for target dialogue data, yet persona-grounded generators optimize individual conversations rather than corpus compositi…

cs.AI2026

PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

Swastik Roy, Rajkumar Pujari, Tharindu Kumarage +5

LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be `…

cs.AI2026

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

Andrea Wynn, Metod Jazbec, Charith Peris +4

Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled desig…

cs.CL2026

SWAN: Semantic Watermarking with Abstract Meaning Representation

Ziping Ye, Gourab Dey, Christos Christodoulopoulos +7

We introduce SWAN (Semantic Watermarking with Abstract Meaning Representation), a novel framework that embeds watermark signatures into the semantic structure of a sentence using A…