12 papers
Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction
Sophia Koehler, Antonia Wüst, Inga Ibs +5
A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments…
COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
David Steinmann, Antonia Wüst, Kristian Kersting +1
While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to s…
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
Berkant Turan, Suhrab Asadulla, David Steinmann +3
While Prover-Verifier Games (PVGs) offer a promising path toward verifiability in nonlinear classification models, they have not yet been applied to complex inputs such as high-dim…
SLR: Automated Synthesis for Scalable Logical Reasoning
Lukas Helff, Ahmad Omar, Felix Friedrich +7
We introduce SLR, an end-to-end framework for systematic evaluation and training of Large Language Models (LLMs) via Scalable Logical Reasoning. Given a user's task specification,…
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Lukas Helff, Quentin Delfosse, David Steinmann +6
As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode emerges: LLMs gaming verifi…
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
Nina Żukowska, Wolfgang Stammer, Bernt Schiele +1
Transparency of neural networks' internal reasoning is at the heart of interpretability research, adding to trust, safety, and understanding of these models. The field of mechanist…