4 papers · 1 filter
Interpretable Reward Modeling with Active Concept Bottlenecks
Sonia Laguna, Katarzyna Kobalczyk, Julia E. Vogt +1
We introduce Concept Bottleneck Reward Models (CB-RM), a reward modeling framework that enables interpretable preference learning through selective concept annotation. Unlike stand…
Beyond Glucose-Only Assessment: Advancing Nocturnal Hypoglycemia Prediction in Children with Type 1 Diabetes
Marco Voegeli, Sonia Laguna, Heike Leutheuser +3
The dead-in-bed syndrome describes the sudden and unexplained death of young individuals with Type 1 Diabetes (T1D) without prior long-term complications. One leading hypothesis at…
Measuring Leakage in Concept-Based Methods: An Information Theoretic Approach
Mikael Makonnen, Moritz Vandenhirtz, Sonia Laguna +1
Concept Bottleneck Models (CBMs) aim to enhance interpretability by structuring predictions around human-understandable concepts. However, unintended information leakage, where pre…
Exploiting Interpretable Capabilities with Concept-Enhanced Diffusion and Prototype Networks
Alba Carballo-Castro, Sonia Laguna, Moritz Vandenhirtz +1
Concept-based machine learning methods have increasingly gained importance due to the growing interest in making neural networks interpretable. However, concept annotations are gen…