- Arizona State UniversityUS1 paper
- Australian National UniversityAU1 paper
- Film IndependentUS1 paper
- Infectious Diseases Research CollaborationUG1 paper
- Lawrence Berkeley National LaboratoryUS1 paper
- Leiden Observatory1 paper
- Leiden UniversityNL1 paper
- Los Alamos National LaboratoryUS1 paper
- Michigan State UniversityUS1 paper
- Stony Brook UniversityUS1 paper
- University of ArizonaUS1 paper
- University of Wisconsin–MadisonUS1 paper
Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
Roland Pihlakas, Jan Llenzl Dagohoy
Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behavio…
cs.CY2026
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
Roland Pihlakas, Sruthi Susan Kuriakose
Many AI alignment discussions of "runaway optimisation" focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective (e.g., "paperclip maximiser", specific…