Showing 2026Show all
3 papers · 1 filter
cs.LG2026
Real-Time Progress Prediction in Reasoning Language Models
Hans Peter Lyngsøe Raaschou-Jensen, Constanza Fierro, Anders Søgaard
Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic tasks. However, as these models operat…
cs.CL2026
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
cs.CY2026
Brainrot: Deskilling and Addiction are Overlooked AI Risks
Ilias Chalkidis, Anders Søgaard
The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b)…