4 papers
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction f…
From Rights to Rites: Expectations Management in Smart-Home AI
Varad Vishwarupe, Ivan Flechais, Marina Jirotka +1
Domestic voice assistants and smart-home devices are increasingly embedded in everyday routines, yet their ethics are often treated as an afterthought or delegated to compliance te…
The Collaboration Gap in Human-AI Work
Varad Vishwarupe, Marina Jirotka, Nigel Shadbolt +1
LLMs are increasingly presented as collaborators in programming, design, writing, and analysis. Yet the practical experience of working with them often falls short of this promise.…
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
Varad Vishwarupe, Ivan Flechais, Nigel Shadbolt +1
Large language models (LLMs) are increasingly integrated into design and development workflows, yet decisions about their use are rarely binary or purely technical. We report findi…