107 citations · 166 across the 26 of their papers we have counts for
8 papers · 1 filter
Position: Theory of Mind Benchmarks are Broken for Large Language Models
Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf +4
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This…
Evaluating the Prompt Steerability of Large Language Models
Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy +5
Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to ev…
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
Shun Ide, Allison Blunt, Djallel Bouneffouf
We propose a general approach to quantitatively assessing the risk and vulnerability of artificial intelligence (AI) systems to biased decisions. The guiding principle of the propo…
Conversational Topic Recommendation in Counseling and Psychotherapy with Decision Transformer and Large Language Models
Aylin Gunal, Baihan Lin, Djallel Bouneffouf
Given the increasing demand for mental health assistance, artificial intelligence (AI), particularly large language models (LLMs), may be valuable for integration into automated cl…
Contextual Moral Value Alignment Through Context-Based Aggregation
Pierre Dognin, Jesus Rios, Ronny Luss +7
Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capabil…
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16
The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…