3 papers
cs.AI2026
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Bo Liu, Leon Guertler, Simon Yu +9
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…
cs.CL2026
Linear Personality Probing and Steering in LLMs: A Big Five Study
Michel Frising, Daniel Balcells
Large language models (LLMs) exhibit distinct and consistent personalities that greatly impact trust and engagement. While this means that personality frameworks would be highly va…
cs.LG2024
Evolution of SAE Features Across Layers in LLMs
Daniel Balcells, Benjamin Lerner, Michael Oesterle +2
Sparse Autoencoders for transformer-based language models are typically defined independently per layer. In this work we analyze statistical relationships between features in adjac…