12 papers
Can LLM Code Explanations Adapt to Diverse Problem-Solvers' Needs?
Andrew Anderson, David Piorkowski, Justin Weisz +2
Large language model (LLM) code explanations can support people in solving code-related problems, yet prior work has shown that people have diverse problem-solving styles. If expla…
AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
Ivoline C. Ngong, Keerthiram Murugesan, Swanand Kadhe +3
Agentic systems are increasingly acting on users' behalf, accessing calendars, email, and personal files to complete everyday tasks. Privacy evaluation for these systems has focuse…
Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering
Justin D. Weisz, Michael Muller, Kush R. Varshney
What better way to understand the impact of AI on software engineering than to ask AI itself? We constructed Story Arena, a multi-agent "writer's room" in which multiple AI agents,…
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
Hyo Jin Do, Rachel Ostrand, Werner Geyer +3
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advan…
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
Ivoline Ngong, Swanand Kadhe, Hao Wang +4
Conversational agents are increasingly woven into individuals' personal lives, yet users often underestimate the privacy risks associated with them. The moment users share informat…
Position: Theory of Mind Benchmarks are Broken for Large Language Models
Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf +4
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This…