5 papers
Real-Time Progress Prediction in Reasoning Language Models
Hans Peter Lyngsøe Raaschou-Jensen, Constanza Fierro, Anders Søgaard
Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic tasks. However, as these models operat…
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
Brainrot: Deskilling and Addiction are Overlooked AI Risks
Ilias Chalkidis, Anders Søgaard
The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b)…
Lost at the Beginning of Reasoning
Baohao Liao, Xinyi Chen, Sara Rajaee +5
Recent advancements in large language models (LLMs) have significantly advanced complex reasoning capabilities, particularly through extended chain-of-thought (CoT) reasoning that…
Federated learning, ethics, and the double black box problem in medical AI
Joshua Hatherley, Anders Søgaard, Angela Ballantyne +1
Federated learning (FL) is a machine learning approach that allows multiple devices or institutions to collaboratively train a model without sharing their local data with a third-p…