12 papers · 1 filter
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina +2
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced d…
Training Language Models to Use Prolog as a Tool
Niklas Mellgren, Peter Schneider-Kamp, Lukas Galke Poech
Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use Prolog as an external symbolic re…
PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
Gianluca Barmina, Federico Torrielli, Sven Harms +7
Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpf…
LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs
Gianluca Barmina, Peter Schneider-Kamp, Lukas Galke Poech
Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under…
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
Stine Lyngsø Beltoft, William Brach, Federico Torrielli +5
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding huma…
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
Federico Torrielli, Peter Schneider-Kamp, Lukas Galke Poech
An activation oracle is a language model trained to read another model's internal activations and describe them in natural language, for example to name a secret word the other mod…