2 papers
cs.LG2026
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
Sushant Gautam, Finn Schwall, Annika Willoch Olstad +6
Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this sett…
cs.SI2026
The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity
Sushant Gautam, Annika W. Olstad, Klas H. Pettersen +1
Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory Archive, an incremental datase…