3 papers
cs.AI2026
The Case for Model Science: Verify, Explore, Steer, Refine
Przemyslaw Biecek, Luca Longo, Jianlong Zhou +3
We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model…
cs.CL2026
Judge Circuits
Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10
LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…
stat.ML2026
-TCAV: A Unified Framework for Testing with Concept Activation Vectors
Ekkehard Schnoor, Jawher Said, Malik Tiomoko +2
Concept Activation Vectors (CAVs) are a fundamental tool for concept-based explainability in deep learning, yet their practical utility is limited by statistical instability. We an…