7 papers · 1 filter
Contextual Value Alignment via Multilayer Combinatorial Fusion
Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants ha…
Mitigating Misalignment Contagion by Steering with Implicit Traits
Maria Chang, Ronny Luss, Miao Liu +3
Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research…
Survey: Multi-Armed Bandits Meet Large Language Models
Djallel Bouneffouf, Raphael Feraud
Bandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-maki…
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
Djallel Bouneffouf, Matthew Riemer, Kush Varshney
This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by huma…
Position: Theory of Mind Benchmarks are Broken for Large Language Models
Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf +4
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This…
Proceedings of 1st Workshop on Advancing Artificial Intelligence through Theory of Mind
Mouad Abrini, Omri Abend, Dina Acklin +105
This volume includes a selection of papers presented at the Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2025 in Philadelphia US on 3rd March 2…