6 papers
Influence of Prompt Engineering on Small Language Models for Guarded Query Routing
Richard Šléher, William Brach, Kristián Košťál +1
We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in-distribution queries or rejec…
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
Stine Lyngsø Beltoft, William Brach, Federico Torrielli +5
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding huma…
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
William Brach, Federico Torrielli, Stine Lyngsø Beltoft +3
Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of…
ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation
William Brach, Francesco Zuppichini, Marco Vinciguerra +1
Producing output that conforms to a specified JSON schema underlies tool use, structured extraction, and knowledge base construction in modern large language models. Despite this c…
SommBench: Assessing Sommelier Expertise of Language Models
William Brach, Tomas Bedej, Jacob Nielsen +10
With the rapid advances of large language models, it becomes increasingly important to systematically evaluate their multilingual and multicultural capabilities. Previous cultural…
Chain of Summaries: Summarization Through Iterative Questioning
William Brach, Kristián Košťál, Lukas Galke Poech
Large Language Models (LLMs) are increasingly using external web content. However, much of this content is not easily digestible by LLMs due to LLM-unfriendly formats and limitatio…