3 papers
cs.AI2025
The Steganographic Potentials of Language Models
Artem Karpov, Tinuade Adeleke, Seong Hah Cho +1
The potential for large language models (LLMs) to hide messages within plain text (steganography) poses a challenge to detection and thwarting of unaligned AI agents, and undermine…
cs.AI2025
Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
Kenneth J. K. Ong, Lye Jia Jun, Hieu Minh "Jord" Nguyen +2
As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperati…
cs.AI2024
Inducing Human-like Biases in Moral Reasoning Language Models
Artem Karpov, Seong Hah Cho, Austin Meek +3
In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same…