Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
CounterMoral: Editing Morals in Language Models
Michael Ripa, Jim Davies
Recent advancements in language model technology have significantly enhanced the ability to edit factual information. Yet, the modification of moral judgments, a crucial aspect of…
cs.AI2026
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…