2 papers
cs.CV2026
BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding
Patrick Knab, Orgest Xhelili, Inis Buzi +7
Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physic…
cs.CV2026
Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
Patrick Knab, Sascha Marton, Philipp J. Schubert +2
Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video rem…