1 citations · 1 across the 3 of their papers we have counts for
7 papers
COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
David Steinmann, Antonia Wüst, Kristian Kersting +1
While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to s…
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
Nils Grandien, David Steinmann, Felix Friedrich +1
Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable,…
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
Berkant Turan, Suhrab Asadulla, David Steinmann +3
While Prover-Verifier Games (PVGs) offer a promising path toward verifiability in nonlinear classification models, they have not yet been applied to complex inputs such as high-dim…
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Lukas Helff, Quentin Delfosse, David Steinmann +6
As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode emerges: LLMs gaming verifi…
Object Centric Concept Bottlenecks
David Steinmann, Wolfgang Stammer, Antonia Wüst +1
Developing high-performing, yet interpretable models remains a critical challenge in modern AI. Concept-based models (CBMs) attempt to address this by extracting human-understandab…
Right on Time: Revising Time Series Models by Constraining their Explanations
Maurice Kraus, David Steinmann, Antonia Wüst +2
Deep time series models often suffer from reliability issues due to their tendency to rely on spurious correlations, leading to incorrect predictions. To mitigate such shortcuts an…