7 citations · 8 across the 6 of their papers we have counts for
8 papers · 1 filter
CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models
Souptik Kumar Majumdar, Fabian Kögel, Andreas Bulling
Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. Howev…
VDial: Unification of Video and Visual Dialog via Multimodal Experts
Adnen Abdessaied, Anna Rohrbach, Marcus Rohrbach +1
We present VDial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multim…
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
Lei Shi, Andreas Bulling
We propose CLAD, a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predictin…
Multi-Modal Video Dialog State Tracking in the Wild
Adnen Abdessaied, Lei Shi, Andreas Bulling
We present MST-MIXER - a novel video dialog model operating over a generic multi-modal state tracking scheme. Current models that claim to perform multi-modal state tracking fall s…
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
Anna Penzkofer, Lei Shi, Andreas Bulling
While Vector Symbolic Architectures (VSAs) are promising for modelling spatial cognition, their application is currently limited to artificially generated images and simple spatial…
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
Lei Shi, Paul Bürkner, Andreas Bulling
We present ActionDiffusion -- a novel diffusion model for procedure planning in instructional videos that is the first to take temporal inter-dependencies between actions into acco…