3 papers
cs.AI2026
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
Matthew Riemer, Tommaso Tosato, Amin Memarian +4
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral cer…
cs.AI2026
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Niklas Herbster, Martin Zborowski, Alberto Tosato +2
Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.…
cs.CL2025
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos +5
Large language models require consistent behavioral patterns for safe deployment, yet there are indications of large variability that may lead to an instable expression of personal…