7 papers
When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
Jerick Shi, Terry Jingcheng Zhang, Bernhard Schölkopf +2
As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions w…
Learning POMDP World Models from Observations with Language-Model Priors
Valentin Six, Frederik Panse, Mathis Fajeau +7
Whether navigating a building, operating a robot, or playing a game, an agent that acts effectively in an environment must first learn an internal model of how that environment wor…
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
Tim Z. Xiao, Johannes Zenn, Zhen Liu +3
Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to generate faithful samples from them. This mi…
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
Florent Draye, Abir Harrasse, Vedant Palit +8
Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders…
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
Egor Shulgin, Dimitri von Rütte, Tianyue H. Zhang +3
Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, wi…
International AI Safety Report 2026
Yoshua Bengio, Stephen Clare, Carina Prunkl +89
The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series…