4 papers
BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding
Patrick Knab, Orgest Xhelili, Inis Buzi +7
Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physic…
Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
Patrick Knab, Sascha Marton, Philipp J. Schubert +2
Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video rem…
From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media
Katharina Prasse, Steffen Jung, Isaac Bravo +4
Social media platforms have become primary arenas for climate communication, generating millions of images and posts that - if systematically analysed - can reveal which communicat…
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
Jannik Brinkmann, Chris Wendler, Christian Bartelt +1
Human bilinguals often use similar brain regions to process multiple languages, depending on when they learned their second language and their proficiency. In large language models…