13 papers
Contextual Value Alignment via Multilayer Combinatorial Fusion
Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants ha…
Mitigating Misalignment Contagion by Steering with Implicit Traits
Maria Chang, Ronny Luss, Miao Liu +3
Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research…
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Aligning large language models (LLMs) with human values is a central challenge for ensuring trustworthy and safe deployment. While existing methods such as Reinforcement Learning f…
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
Matthew Riemer, Erik Miehling, Miao Liu +2
Although parameter-efficient fine-tuning methods, such as LoRA, only modify a small subset of parameters, they can have a significant impact on the model. Our instruction-tuning ex…
Survey: Multi-Armed Bandits Meet Large Language Models
Djallel Bouneffouf, Raphael Feraud
Bandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-maki…
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
Djallel Bouneffouf, Matthew Riemer, Kush Varshney
This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by huma…